HOYALABA PERSONAL RESEARCH NOTEBOOK

Paper Readings / 2026-08-10 / 22 MIN

Reading Diffusion Policy, part 2: why it trains stably, and what the robots show

Sections 4 to 9 of Diffusion Policy: modes and basins, why learning the score removes the normalizing constant, the benchmark and real-robot results, and the limits the authors state.

This post continues from part 1 and covers the second half of Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Part 1 covered Sections 1 to 3: the formulation of a robot policy as a conditional denoising diffusion process, and the network design. This time the scope is Section 4 (Intriguing Properties), Section 5 (Evaluation), Sections 6 and 7 (real-robot experiments), Section 8 (Related Work), and Section 9 (Limitations).

In part 1, I admitted I'd understood maybe half of the energy discussion: why learning an energy makes training unstable, and how Diffusion Policy avoids that. The question is:

If learning the energy is unstable, why is learning the score stable?

The paper's answer is in Section 4.4, and this time I'll go through it properly.

Before that, Section 4.1 takes the multimodality story from part 1 one step further, using words like mode and basin without explaining them. I spent a while pinning those down, so I've written that part out carefully. Let's go slowly.

Section 4.1: where does multimodality come from?

Section 4 is "Intriguing Properties of Diffusion Policy": property by property, why this policy representation works well. The first is multimodal action distributions. The paper says, briefly, that Diffusion Policy's multimodality comes from two places:

  • Stochastic initialization: every sample starts from a fresh AtKA_t^K drawn from a standard normal.
  • Stochastic sampling: Gaussian noise is added at every denoising iteration.

Together, it says, these "specify different convergence basins" and let samples "move between multiple multimodal action basins." Short sentences, but I got stuck here for a long time, because I didn't have a precise idea of what a mode or a basin was. So for anyone who trips on the same spot, here's what I learned while looking it up.

Wait: what exactly is a mode?

In statistics, the mode is the most frequent value: the point of highest probability density. (The word apparently comes from the sense of "what's in fashion": the most common value is the fashionable one. The modality that confused me in part 1, meaning a form of data, grew from the same root meaning "manner" or "form," so one word carrying two meanings isn't completely arbitrary.)

In action space, though, it helps to think of a mode as a cluster rather than a single value. Back to the T-block: all the actions that amount to "go around the left" form one cluster, and all the "go around the right" actions form another. There are many slightly different ways to go around the left, but they all sit inside that cluster.

Mode: the cluster of actions that belong to one strategy, and the point of highest probability within it.

Seen this way, "two modes" doesn't mean "exactly two correct actions." It means "two strategies."

Peak or valley?

This is where I really got stuck. In part 1, I described a multimodal distribution as one with several peaks. But here the paper uses basin, a word for a hollow. It first called them peaks, and now suddenly they're pits? The two seemed to contradict each other.

The cause was that they look at different axes. Recall the energy-based policy from part 1:

p(a∣o) ∝ e−E(o,a)p(a \mid o)\ \propto\ e^{-E(o,a)}

There's a minus sign in the exponent, so the lower the energy, the higher the probability. The spots with high probability and the spots with low energy are exactly the same spots.

  • Plot probability on the vertical axis, and those spots stick up: peaks.
  • Plot energy on the vertical axis, and the same spots sink down: valleys.

With that in mind, Equation (2) from part 1 looks different too:

x′=x−γ∇E(x)x'=x-\gamma\nabla E(x)

One denoising step moves a little downhill in energy. So generating a sample with diffusion is walking down into a valley of this landscape: starting from a lump of noise and rolling, step by step, toward high probability.

Then what is a basin?

In geography, a basin is a watershed: the area of land that drains into one river. The meaning here is almost the same.

Basin: the set of starting points that all end up in the same valley if you roll downhill.

Picture a map with two valleys and a ridge between them. A raindrop that falls left of the ridge flows into the left valley; one that falls to the right flows into the right valley. All the land that drains into the left valley is the left valley's basin. So:

  • Mode: the destination, the bottom of the valley itself.
  • Basin: the region of starting points that lead to that destination.

If you know optimization, you can think of a basin as the region around an optimum, including the bowl-shaped surroundings that lead into it.

Back to the paper: randomness in two places

First, stochastic initialization. Diffusion Policy starts every sample from a new AtKA_t^K drawn from a standard normal. In the map analogy, that's choosing afresh where the raindrop lands. Which basin it lands in largely decides which valley it flows to. That's what the paper means when it says the initial sample helps specify different convergence basins. So even with the same observation OtO_t, one run might go around the left and the next around the right, because the starting point differs each time.

Second, stochastic sampling. In part 1's Equation (4), Gaussian noise is added back at every denoising step. That's like shaking the raindrop a little on its way down. With luck, it can hop a low ridge into the neighboring basin. That's what the paper means by samples moving between action basins.

At first I didn't see why the second part was needed. Wouldn't a random starting point be enough? But without the shaking, the drop would be trapped in whichever basin it first landed in, with no way out if the starting point fell somewhere awkward.

So why don't the actions jitter?

A natural question follows, and it stopped me again. If samples shake at every step and can hop between basins, won't the robot also swing left, then right, then left?

It doesn't. And this is, for me, the most important point in Section 4.1. The shaking happens inside the denoising process that produces one action sequence. After KK iterations, the result is a single At0A_t^0 that has already settled into one valley. The robot executes that final result, not the shaking along the way. Recall the kk and tt distinction from part 1:

The shaking during denoising happens along the kk axis. Jittery behavior would show up along the tt axis.

And there's one more layer. Diffusion Policy generates an entire action sequence at once. One round of denoising decides, as a single piece, how the robot will move over the next several steps. So left and right can't get mixed within a sequence.

Being able to represent several modes and committing to one mode for a whole execution are two different things.

The paper's Figure 3 shows exactly this contrast. In a state where the end-effector could go either left or right around the block, each method behaves differently:

  • Diffusion Policy learns both modes and, within one rollout, commits to one side all the way.
  • LSTM-GMM and IBC are biased toward one mode. They don't properly learn that there are two options.
  • BET captures the modes but fails to commit. Without temporal consistency, it picks a different mode at every step.

The third struck me most. Knowing the options isn't enough; committing to one of them consistently is a separate ability.

Section 4.2: position control works better?

This section surprised me. The paper reports that Diffusion Policy consistently did better with a position-control action space than with velocity control. The authors themselves call it a surprising result, since most recent behavior-cloning work uses velocity control. They offer two guesses:

  • Action multimodality is more pronounced under position control than velocity control. Diffusion Policy represents multimodality well, so this hurts it less than other methods.
  • Position control is less prone to compounding error than velocity control, so it suits action-sequence prediction better.

So if other methods used velocity control because they couldn't cope with multimodality, Diffusion Policy escapes that constraint and gets the benefits of position control.

That made me ask why velocity control had been the better fit all along, and the authors had just answered it. Earlier methods couldn't produce consistent action sequences or handle multiple modes robustly, so making a sequence of positions consistent was hard. Diffusion Policy can represent several modes while still executing one of them consistently, and that's what makes position control viable.

The paper's Figure 4 shows it: switching from velocity to position control lowers LSTM-GMM and BET, while Diffusion Policy improves.

Section 4.3: what predicting a sequence buys

Part 1 introduced ToT_o, TpT_p, and TaT_a; this section explains why predict a sequence at all.

The authors first note that most policy learning avoids sequence prediction, because sampling from a high-dimensional output space is hard. IBC struggles to sample high-dimensional actions from a bumpy energy landscape, and BC-RNN and BET need the number of modes in the action distribution specified in advance. Having sorted out modes, this made more sense to me: to use a GMM or k-means, someone has to decide up front how many peaks there are. DDPM, by contrast, scales to large outputs without losing expressiveness, as image generation already showed. That buys two things:

  • Temporal action consistency: if each step's action is sampled from its own independent multimodal distribution, consecutive actions can come from different modes, producing jittery back-and-forth motion.
  • Robustness to idle actions: when a demonstrator pauses, the data fills with repeated positions or near-zero velocities. Single-step policies easily overfit to these stretches and stop.

The paper notes that idle actions are common in teleoperation, and even necessary for tasks like pouring liquid. BC-RNN and IBC often got stuck in real experiments unless idle actions were explicitly removed from the training data. This comes back in the real-robot results.

Section 4.4: so why is training stable?

This is the part I understood only half of last time. Let's do it properly.

The paper first restates that an implicit policy represents the action distribution as an energy-based model:

pθ(a∣o)=e−Eθ(o,a)Z(o,θ)p_\theta(a \mid o)=\frac{e^{-E_\theta(o,a)}}{Z(o,\theta)}

That's Equation (6), which I pulled forward in part 1. The problem is Z(o,θ)Z(o,\theta) in the denominator: it requires integrating over all actions aa, so it can't be computed directly. EBMs are therefore trained with an InfoNCE-style loss.

The name InfoNCE appeared out of nowhere, so I stopped again. It stands for Info Noise-Contrastive Estimation, and the idea fits in one line:

InfoNCE: instead of computing the probability directly, mix the true answer with several fakes and turn it into a multiple-choice question: pick the real one.

It changes the question:

  • Original question: "What is the probability of action aa in this situation?"
  • New question: "Which of these options was actually in the demonstration?"

Answering the first needs Z(o,θ)Z(o,\theta); the second is just a classification problem. The fake options are called negative samples. We've seen this family before: the contrastive loss used by models like CLIP, which pairs images with text, works the same way. Here's the paper's version:

LinfoNCE=−log⁡(e−Eθ(o,a)e−Eθ(o,a)+∑j=1Nnege−Eθ(o,a~j))\mathcal{L}_{\mathrm{infoNCE}}=-\log\left(\frac{e^{-E_\theta(o,a)}}{e^{-E_\theta(o,a)}+\sum_{j=1}^{N_{neg}}e^{-E_\theta(o,\tilde{a}_j)}}\right)

This is Equation (7). It looks complicated, but it does something simple. The numerator holds the true action from the demonstration; the denominator holds the true action plus the fake actions a~j\tilde{a}_j. To make this large, lower the energy of the true action and raise the energy of the fakes. The bundle of fake samples in the denominator effectively stands in for Z(o,θ)Z(o,\theta).

And there's the problem. The estimate depends heavily on how many fakes you draw and where you draw them from, and when the estimate is off, training wobbles. That's what "unstable because of negative sampling" meant in part 1.

Now, how Diffusion Policy avoids it:

∇alog⁡p(a∣o)=−∇aEθ(a,o)−∇alog⁡Z(o,θ)⏟= 0≈−ϵθ(a,o)\begin{aligned} \nabla_a \log p(a \mid o) &= -\nabla_a E_\theta(a,o)-\underbrace{\nabla_a \log Z(o,\theta)}_{=\,0} \\ &\approx -\epsilon_\theta(a,o) \end{aligned}

This is Equation (8), the most important equation in this post. The key is the underbraced term. Z(o,θ)Z(o,\theta) depends only on the observation oo and the parameters, not on the action aa. Differentiate with respect to aa, and it's simply zero. The ZZ that was unavoidable when working with the distribution itself vanishes entirely once we work with the gradient of the log probability.

The noise-prediction network ϵθ(a,o)\epsilon_\theta(a,o) approximates exactly the negative of this score function. That's why Z(o,θ)Z(o,\theta) never appears in the inference Equation (4) or the training Equation (5) from part 1.

In the landscape picture from Section 4.1: knowing the exact altitude of a valley floor is hard, but knowing which way is downhill from where you stand is much easier. And rolling downhill only needs the second.

The paper's Figure 6 shows this experimentally. For IBC, the training loss on the energy function falls smoothly, yet it fails to fit the training actions, and its evaluation performance keeps oscillating. So IBC-style work evaluates every checkpoint and reports the best one, which on a real robot is a real burden: every checkpoint has to be run on hardware.

Section 4.5: the connection to control theory

This section checks whether, in the simplest case, Diffusion Policy recovers the control law we already know. Assume a linear system and a linear feedback policy:

st+1=Ast+Bat+wt,at=−Ksts_{t+1}=As_t+Ba_t+w_t,\qquad a_t=-Ks_t

Such a policy could come from solving LQR, for example, and of course you don't need diffusion to imitate it. It's still worth checking. With a prediction horizon of 1, we can find the optimal denoiser that minimizes

L=MSE(ϵk, ϵθ(st,−Kst+ϵk,k)).\mathcal{L}=\mathrm{MSE}\big(\epsilon_k,\ \epsilon_\theta(s_t,-Ks_t+\epsilon_k,k)\big).

This is Equation (9). The paper shows the optimal denoiser takes the form ϵθ(s,a,k)=1σk[a+Ks]\epsilon_\theta(s,a,k)=\frac{1}{\sigma_k}[a+Ks], and DDIM sampling converges to a=−Ksa=-Ks. In the simple case, Diffusion Policy lands back on the feedback law we'd expect.

I found the next paragraph even more interesting. With a prediction horizon longer than 1, predicting a future action at+t′a_{t+t'} from the current state sts_t alone requires the optimal denoiser to output at+t′=−K(A−BK)t′sta_{t+t'}=-K(A-BK)^{t'}s_t. The system matrices AA and BB appear in it.

To perfectly clone a state-dependent behavior, the learner has to implicitly learn the dynamics model the task requires.

So predicting a long action sequence isn't just a longer output; the model has to know, to some degree, how the situation changes when it acts. The paper also draws a line at the end: if the plant or the policy is nonlinear, predicting future actions gets much harder and becomes multimodal again. So I read this section not as "Diffusion Policy solved control mathematically," but as "in the simple case, it doesn't contradict the control-theory view."

Section 5: evaluation across 15 tasks and 4 benchmarks

Section 5 evaluates 15 tasks across 4 benchmarks. The simulated environments:

  • Robomimic: five tasks (Lift, Can, Square, Transport, ToolHang), each with proficient-human (PH) and mixed-quality (MH) datasets, for nine variants.
  • Push-T: from IBC. A circular end-effector pushes a T-shaped block to a target pose. Contact-rich and precision-demanding.
  • Multimodal Block Pushing: from BET. Push two blocks into two squares, in either order. The arbitrary order creates long-horizon multimodality.
  • Franka Kitchen: seven objects; each demonstration completes four tasks in an arbitrary order. Short- and long-horizon multimodality at once.

The evaluation protocol is worth noting. The paper uses 3 training seeds and 50 environment initializations, and reports the average of the last 10 checkpoints saved every 50 epochs. To match the original papers, it also reports the best-checkpoint number, so the tables read "max / average of last 10." I liked that: having criticized IBC in Section 4.4 for needing checkpoint selection, the paper reports its own results by the same yardstick.

There's also an honest detail. A footnote says a bug in the evaluation code meant only 22 initial conditions were used for the robomimic tasks, and adds that all baselines were evaluated the same way, so the conclusions don't change. The acknowledgments even thank the person who found the bug.

Section 5.3: key findings

The headline: across every task and variant, with state or vision observations, Diffusion Policy outperforms prior methods, with an average improvement of 46.9%. The paper's key findings come to about six:

1. It represents short-horizon multimodality, which the paper defines as multiple ways of achieving the same immediate goal. The left-or-right T-block choice is exactly this.

2. It handles long-horizon multimodality too: completing different sub-goals in inconsistent orders. There are several goals, and no fixed order. Touching Kitchen's seven objects in any order, or pushing either block first in Block Push. Same word, different grain: a branch in the moment versus a branch in the overall order. The numbers differ a lot here: 32% better on Block Push's p2 metric, and 213% better on Kitchen's p4 metric. (pxp_x means, for Block Push, the frequency of pushing xx blocks into targets; for Kitchen, the frequency of interacting with at least xx objects. So p4p_4 is the hardest metric, and that's where the gap is widest.)

3. It makes better use of position control, confirming Section 4.2 experimentally.

4. There's a trade-off in the action horizon. A horizon above 1 helps action consistency and idle-action handling, but too long a horizon slows reactions and hurts performance. The paper confirms this experimentally, with 8 steps best on most tasks tested.

5. It's robust to latency. Thanks to receding-horizon position control, it kept peak performance with up to 4 steps of latency. That matters on real robots, where image processing, policy inference, and network delays all exist.

6. Training is stable. The best hyperparameters were largely consistent across tasks.

(The paper's Figure 5 shows the action-horizon and latency ablations.)

Section 5.4: how should the vision encoder be used?

This ablation was added in the extended version, and it's probably what practitioners most want to know. On robomimic's Square task, the paper compares three architectures (ResNet-18, ResNet-34, ViT-B/16) under three training strategies (from scratch, frozen pretrained encoder, fine-tuned pretrained encoder). Three results stand out:

  • Training a ViT from scratch doesn't work well: 22% success, which the paper attributes to limited data.
  • Freezing a pretrained encoder performs poorly: for ResNet-18, success drops from 0.94 to 0.58.
  • Fine-tuning a pretrained encoder with a small learning rate works best: a CLIP-trained ViT-B/16 reaches 98% in 50 epochs.

The paper's interpretation of the second result is interesting: the visual representation Diffusion Policy wants differs from what generic pretraining provides. So you can't just plug a pretrained encoder in; it has to move at least a little along with the policy. The paper also notes the final gaps between architectures weren't as large as their theoretical capacity differences, and cautiously adds that the gap may widen on more complex tasks.

Section 6: the real robot, Push-T

From Section 6 on, the paper runs real robots, and I liked how much of the paper this takes up. The real Push-T is harder than the simulated one, for three reasons:

  • It's multi-stage. After pushing the T into place, the end-effector has to retreat to a designated zone so it doesn't block the camera.
  • Fine adjustments are needed to seat the block fully in the target, which adds multimodality at the stage transitions.
  • IoU is measured at the final step, not as the maximum over all steps.

The setup is a UR5. The policy commands at 10 Hz, linearly interpolated to 125 Hz for execution. The reported results:

  • Diffusion Policy (end-to-end): 95% success, average IoU 0.80
  • Human: 100% success, average IoU 0.84
  • Best LSTM-GMM variant: 20%
  • Best IBC variant: 0%

The failure modes line up with the theory. The baselines mostly break down at stage transitions, where multimodality is strong and decision boundaries are fuzzy. LSTM-GMM got stuck near the T block in 8 of 20 trials; IBC ended the pushing stage too early in 6 of 20.

One condition the paper adds is important: because of the task requirements, idle actions weren't removed from the training data, and that contributed to LSTM and IBC overfitting to small actions and stopping. The idle-action problem from Section 4.3, showing up as real numbers.

The pretrained-encoder comparison appears here too. R3M reaches a respectable 80%, but its actions are jittery and it often stalls; ImageNet pretraining is much worse at 15%. The best models were all trained end-to-end.

Finally, the perturbation tests:

  • Covering the front camera by hand for 3 seconds: the policy jitters a little but keeps its path and pushes the block in.
  • Moving the block during fine adjustment: it immediately replans to push from the other side.
  • Moving the block after stage one, while the arm heads to the end zone: it turns back, realigns the block, then goes to the end zone.

The paper states that the last behavior never appeared in the demonstrations, and cautiously suggests the policy may be able to synthesize new behavior for unseen observations. The "may be able to" suggests the authors don't want to overclaim here either.

Sections 6.2 and 6.3: mug flipping and sauce

The other single-arm tasks each target something specific.

Mug flipping tests complex 3D rotations and motion near the hardware's kinematic limits. Depending on the mug's initial pose, the demonstrator either places it in the right orientation directly or pushes the handle once more to rotate it, and grasps it in various ways. In the vocabulary from earlier, the data doesn't just have several modes; the kinds of modes vary. The result: 90% success over 20 trials. LSTM-GMM, trained on part of the same data, failed to grasp in all 20. The paper also notes behaviors that weren't in the demonstrations, such as pushing the handle several times to align it and re-grasping a dropped mug.

Pouring and spreading sauce tests non-rigid objects, 6-DoF actions, and periodic motion. Pouring requires the robot to wait in place for a while as viscous sauce fills the ladle: exactly the idle-action problem. Spreading follows a human chef's spiral, needing long-period repetitive patterns and short-horizon feedback at once. The reported results: pouring at 0.79 success (human 1.00, LSTM-GMM 0.00) and spreading at 1.00 (human 1.00, LSTM-GMM 0.00). Both tasks used the same hyperparameters as Push-T and succeeded on the first attempt. That last sentence impressed me more than the numbers.

Section 7: bimanual tasks, and a story about haptics

This section was added in the extended version. It extends the observation and action spaces to two arms and runs three tasks. The paper says most of the effort went into the robot stack, supporting multi-arm teleoperation and control, and that Diffusion Policy itself worked without any hyperparameter tuning. The results:

  • Egg beater: 55% success over 20 trials, 210 demonstrations
  • Mat unrolling: 75% over 20 trials, 162 demonstrations
  • Shirt folding: 75% over 20 trials, 284 demonstrations

What stayed with me from this section, though, was the data collection. For the egg-beater task, teleoperating without haptic feedback, even an expert succeeded zero times in 10 attempts: 5 times they pulled the crank handle off, 3 times they lost the handle, and 2 times they hit the torque limit. With haptic feedback, the same person succeeded 10 out of 10.

Before the policy, the first bottleneck may be whether you can collect good demonstrations at all.

When reading behavior-cloning papers, my eye usually goes to the model architecture. This paragraph made me look at the data side again.

Section 8: where the paper stands

Section 8 splits behavior-cloning methods into two families by policy structure, the same split as Figure 1 in part 1.

Explicit policies map observations directly to actions: simple to train with a regression loss, one forward pass at inference. To handle multimodal behavior, some discretize the action space into a classification problem, but the number of bins grows exponentially with dimension. Mixture density networks, or clustering followed by offset prediction, are other options, but the paper notes they're sensitive to hyperparameters and prone to mode collapse. (Mode collapse: the distribution has several peaks, but the model learns only one and misses the rest.)

Implicit policies define the distribution with an EBM. They can give low energy to many actions, which suits multimodality, but negative sampling in the InfoNCE loss makes training unstable.

Then there are diffusion models, and here the paper is clear about its position. It distinguishes itself from work that uses diffusion as a component of planning or reinforcement learning: this paper uses diffusion as the visuomotor control policy itself, trained by behavior cloning. It's candid about concurrent work, too. Others focused on efficient sampling strategies or goal conditioning with classifier-free guidance; this paper focused on effective action spaces, and the simulation conclusions largely agree. Its distinguishing contribution, it says, is real-robot evidence for receding-horizon prediction, the choice between velocity and position control, and the importance of how visual conditioning is done.

Section 9: the limits the paper states itself

The paper closes with two limitations.

The first is that it inherits the limits of behavior cloning. With too little or suboptimal demonstration data, performance is capped. In part 1, distinguishing the energy function from a value function, I said "this paper has no rewards." Here's the price: with no reward, there's no way to judge whether a demonstration is good or bad, and a bad demonstration gets imitated faithfully. The paper suggests combining it with paradigms like reinforcement learning that can also use suboptimal or negative data.

The second is computational cost and inference latency. It's more expensive than simple methods like LSTM-GMM. Predicting action sequences partly mitigates this, but the paper says it may not be enough for tasks needing high control rates. It points to diffusion-acceleration work, such as new noise schedules, inference solvers, and consistency models, as future directions. In part 1 we saw DDIM cut inference to 10 steps for 0.1 s; the paper itself doesn't claim that's enough.

What I took away

Having read it across two posts, I'd sum up what I want to keep in three points.

First, diffusion isn't an image model. At heart, diffusion is a way to draw samples from a complex distribution. For images, the sample happened to be an image; here, it's an action sequence.

Second, learning the score makes the uncomputable constant disappear. The single underbraced zero in Equation (8) holds up the paper's entire stability claim. When I first met the score-based view in the previous paper, I didn't really feel why it was better. Here, it shows up as a very concrete benefit. And the landscape analogy from Section 4.1 applies again: knowing which way is downhill is far easier than knowing a valley's absolute altitude, and it's all you need.

Third, architecture alone doesn't make it work on a robot. Receding horizons, position control, FiLM conditioning, an end-to-end vision encoder, DDIM acceleration, and even haptic feedback: more than half of the paper is really about these practical choices.

Summary

  • I started with what modes and basins are: peaks in probability are valleys in energy, and a basin is the region of starting points that drain into a valley.
  • Diffusion Policy draws a fresh starting point each time and shakes samples on the way down, so it can represent several modes. But because it generates a whole action sequence at once, each execution commits to a single mode.
  • The heart of this post is Equation (8). Learning the gradient of the log probability makes the intractable normalizing constant Z(o,θ)Z(o,\theta) vanish under differentiation, so training needs no negative sampling. That's what the paper's training stability really rests on, and what separates it from IBC.
  • In experiments, the paper reports an average 46.9% improvement across 15 tasks and 4 benchmarks, with the biggest gap on Kitchen's p4 metric, which needs long-horizon multimodality.
  • On real robots it approaches human-level results (Push-T 95%, mug flipping 90%) and extends to bimanual tasks without hyperparameter tuning.
  • The paper is upfront about the limits of behavior cloning and its inference cost, and leaves those to future work.

Next, I'd like to look at this paper from the implementation side: what input and output shapes a real Diffusion Policy is trained with.