VLAs such as π0 generate robot actions by "starting from noise and adding velocity a few times." This post follows that process in the same order as my handwritten study notes. First we walk through the three steps a robot takes to produce an action, and then we turn to the question: why predict a velocity at all?
How to read it: yellow boxes are my original notes, and green boxes refine them into something more precise. Proofs and derivations are folded into Go deeper boxes, so if you only want the main thread you can skip them and the text still flows. Every figure is computed live in your browser, so try moving the sliders as you read.
0The big picture: letting noise flow into actions
Look at the figure at the top again. Points bunched in the middle (Gaussian noise) flow outward and arrive at eight clusters (the data). At every instant, each point follows an arrow telling it which direction to move from where it currently is, and how fast. The collection of all these arrows is called a velocity field \(v(x,t)\).
In one sentence, flow matching is a way to learn this velocity field with a neural network. Once you know the field, you can generate data by starting from noise and moving along the arrows a little at a time. For a robot, the "data" is the trajectory of actions it is about to execute.
Throughout this post, time is written \(t\) or \(\tau\), with 0 being noise and 1 being data (actions). In the VLA sections I use \(\tau\), as the π0 paper does, to avoid confusion with the robot's control time step.
Go deeper The precise statement with ODEs and the continuity equation
Moving along a velocity field is an ordinary differential equation (ODE):
$$\frac{d}{dt}\,\psi_t(x) = v_t\big(\psi_t(x)\big),\qquad \psi_0(x)=x.$$\(\psi_t\) is called the flow. The distribution obtained by transporting the noise distribution \(p_0=\mathcal N(0,I)\) along the flow is written \(p_t=[\psi_t]_\#\,p_0\), and the goal is for \(p_1\) to equal the data distribution \(q\).
The velocity field and the evolving distribution are tied together by the continuity equation, since probability mass is neither created nor destroyed but only moves along the velocity:
$$\frac{\partial p_t(x)}{\partial t} + \nabla\cdot\big(p_t(x)\,v_t(x)\big) = 0.$$When this holds we say "\(v_t\) generates \(p_t\)." A model whose velocity field is a neural network is called a continuous normalizing flow (CNF). CNFs were originally trained by maximum likelihood, which requires solving the ODE during training and is expensive. Flow matching trains them with plain regression and no simulation.
1Step 1. Start from noise
Step 1. Sample the initial noise at \(\tau=0\). Sample Gaussian noise \(\epsilon\sim\mathcal N(0,I)\) that carries no meaningful information. This noise becomes the starting point of inference, \(a_0=\epsilon\).
Exactly as written. One addition: the robot does not produce a single action for one instant, but a bundle of actions for the next \(H\) steps. This is called an action chunk, and in π0, \(H=50\).
$$\mathbf A_t=[\mathbf a_t,\mathbf a_{t+1},\dots,\mathbf a_{t+H-1}]\in\mathbb R^{H\times d_a}$$So the starting noise is a matrix of the same shape, \(\mathbf A^{0}=\epsilon\in\mathbb R^{H\times d_a}\). Think of it as a trajectory whose 50 waypoints are scattered at random.
2Step 2. Read the observation once with the VLM
Step 2. Compute VLM backbone features and cache the KV (run once). o: camera images, ℓ: language instruction → VLM backbone → vision-language features. KV caching: vision-language representation ← masking → action generation (like a one-way street). So the model computes the vision-language key-value (KV) features only once → cached in memory, reused at every later integration step.
π0 has two parts: a large VLM backbone (PaliGemma, about 3B parameters) that understands images and language, and a small action expert (about 300M) dedicated to actions. They have separate weights but share the transformer's attention.
The "one-way street" in my notes is the key idea. The attention mask is set up as below, so action tokens can see the image and language tokens, but image and language tokens cannot see the action tokens.
As a result, the image and language computations come out exactly the same no matter how the actions change. During Step 3 the action is rewritten ten times while this part never changes, so it can be computed once, stored (KV caching), and reused. The heavy 3B model runs once and only the light 300M model runs ten times, which is why it is fast.
The direction of the block matters. If image and language tokens could see action tokens, their computations would change every time the actions change, and reusing the cache would give the wrong answer. Thanks to the mask, caching is not an approximation but exactly the same computation. The robot state \(\mathbf q_t\) is given its own block for the same reason, so it can be cached too.
Go deeper Why caching is exact, and its cost
Split the tokens into three blocks \([\text{images, language}],\ [\mathbf q_t],\ [\mathbf A^\tau]\). Each block attends only to itself and earlier blocks (bidirectionally within a block). Then at layer \(\ell\), the prefix hidden states \(h^\ell_P\) are functions of the prefix inputs alone, and so \(K^\ell_P,V^\ell_P\) are independent of \(\tau\) and \(\mathbf A^\tau\). At each integration step, computing \(Q,K,V\) only for the action tokens and attending over \([K_P;K_A],[V_P;V_A]\) is mathematically identical to recomputing everything.
With prefix length \(L_P\), \(L_A=H\) action tokens and \(N\) integration steps, the pass through the VLM weights happens once (\(O(L_P^2)\) attention), while the action expert runs \(N\) times (\(O(N L_A(L_P+L_A))\)).
Each action token is made by embedding a waypoint with a linear layer, combining it with a sinusoidal embedding of \(\tau\), and passing the result through an MLP. At the output, a linear layer reads a velocity vector from each token. A side benefit: the pretrained VLM representation is not contaminated by noisy action tokens.
3Step 3. Integrate by adding velocity
Step 3. Numerical integration loop (\(\tau=0\to1\), repeated). According to a fixed number of steps \(N\), increase \(\tau\) from 0 to 1 by \(\Delta\tau=1/N\) using Euler's method. ① Feed the current state \(a_{\tau_k}\), the cached KV, and the current time step \(\tau_k\) into the action head ② predict \(v_\theta(a_{\tau_k},o,\ell,\tau_k)\) ③ \(a_{\tau_{k+1}}=a_{\tau_k}+\Delta\tau\cdot v_\theta\) ④ final action trajectory at \(\tau=1\) → action chunk (noise fully removed).
The flow in my notes is correct. Euler's method repeats "move for \(\Delta\tau\) at the current velocity." π0 uses \(N=10\), moving ten times with \(\Delta\tau=0.1\).
- Pass the images, language and state through the VLM once and cache the KV. (Step 2)
- Sample \(\mathbf A^{0}=\epsilon\sim\mathcal N(0,I)\). (Step 1)
- Repeat ten times: the action expert predicts a velocity, and the action moves by \(0.1\times\) that velocity. (Step 3)
- \(\mathbf A^{1}\) is the action chunk. The robot executes the first part of it, then starts again from step 1 with a new observation.
Go deeper Boundary conditions and the code's time convention
The time fed to the model at the last step is \(\tau_{N-1}=1-\Delta\tau\), so the model is never evaluated at \(\tau=1\) during integration. As we will see, the velocity field becomes singular like \(1/(1-\tau)\) as \(\tau\to1\), so this matters.
The public implementation, openpi, uses the opposite time direction from this post: \(t=1\) is noise and \(t=0\) is the action, with \(x_t=t\,\epsilon+(1-t)\mathbf A\), target velocity \(\epsilon-\mathbf A\), and sampling that starts at \(t=1\) and steps with \(dt=-1/N\). Substituting \(t=1-\tau\) and flipping the sign gives exactly the same equations, but mixing formulas from the paper and the code can break everything with a single sign.
4Why predict a velocity?
\(a_\tau=\tau a+(1-\tau)\epsilon\) … linear interpolation. \(\dfrac{da_\tau}{d\tau}=a-\epsilon=u\). When walking straight at constant speed from the noise point \(\epsilon\) to the target \(a\), the direction and velocity to follow is exactly \(u=a-\epsilon\). ∴ Given the current position \(a_\tau\) and the camera/language information, the model \(v_\theta\) is trained to match the target velocity \(u=a-\epsilon\).
Now for how the \(v_\theta\) used in Step 3 is trained. The description in my notes is the training method itself. Take one ground-truth action chunk \(\mathbf A\) from the demonstrations, sample noise \(\epsilon\) and a time \(\tau\), and form a point on the straight line between them. Then reduce the error between the model's predicted velocity at that point and \(\mathbf A-\epsilon\).
No ODE solving, no likelihood computation. It is just regression, and that is why flow matching is easy to train.
But there is a catch
More than one line passes through the same point. Lines from different (noise, target) pairs can pass through the same point at the same time, and they point in different directions. When one point has several correct answers, a model that minimizes squared error ends up outputting the average of those answers.
So the velocity the model actually learns is not the velocity of any individual line but an average velocity. In the figure below, the left panel shows the lines used in training, and the right panel shows the velocity field formed by averaging them, along with trajectories that follow it. Move the slider and you will see that the trajectories on the right bend instead of staying straight, and never cross each other.
Conditional paths used in training straight, crossing
Marginal velocity field and ODE trajectories curved, never crossing
"Walking straight" describes the path of a single pair built during training. What the model learns is the average of many lines, so the trajectory it actually follows at inference time bends. And since the velocity at a given location is a single value, trajectories cannot cross. The lines that used to cross now bend to avoid one another.
Go deeper Why learning the average is fine: the two theorems of conditional flow matching
What we really want is the velocity field \(u_t(x)\) that transports the whole distribution. Regressing it directly gives
$$\mathcal L_{\mathrm{FM}}(\theta)=\mathbb E_{t,\;x\sim p_t}\big\lVert v_\theta(x,t)-u_t(x)\big\rVert^2,$$which cannot be computed since \(u_t\) is unknown. Lipman et al. (2023) first define a path \(p_t(x\mid x_1)\) and velocity \(u_t(x\mid x_1)\) conditioned on a single data sample \(x_1\), and then define the whole by mixing them:
$$p_t(x)=\int p_t(x\mid x_1)q(x_1)\,dx_1,\qquad u_t(x)=\int u_t(x\mid x_1)\,\frac{p_t(x\mid x_1)q(x_1)}{p_t(x)}\,dx_1.$$The fraction is the posterior \(p(x_1\mid x_t=x)\). So \(u_t(x)\) is "a weighted average of the conditional velocities passing through \(x\)," which is precisely the "average velocity" in the main text.
Proof. \(\partial_tp_t(x)=\int\partial_tp_t(x\mid x_1)q\,dx_1=-\int\nabla\cdot(p_t(x\mid x_1)u_t(x\mid x_1))q\,dx_1=-\nabla\cdot(p_t(x)u_t(x))\). The last equality follows from the definition of \(u_t\). \(\square\)
Now write the computable conditional objective:
$$\mathcal L_{\mathrm{CFM}}(\theta)=\mathbb E_{t,\;x_1\sim q,\;x\sim p_t(\cdot\mid x_1)}\big\lVert v_\theta(x,t)-u_t(x\mid x_1)\big\rVert^2.$$Proof. Expanding the squares, the \(\lVert v_\theta\rVert^2\) terms agree because \(\mathbb E_{q(x_1)p_t(x\mid x_1)}=\mathbb E_{p_t(x)}\), and the cross terms agree because \(\mathbb E_{p_t(x)}\langle v_\theta,u_t(x)\rangle=\int\langle v_\theta,\int u_t(x\mid x_1)p_t(x\mid x_1)q\,dx_1\rangle dx=\mathbb E_{q,p_t(x\mid x_1)}\langle v_\theta,u_t(x\mid x_1)\rangle\). The remaining terms do not depend on \(\theta\). \(\square\)
For the linear interpolation in my notes, \(x_t=tx_1+(1-t)x_0\), we get \(p_t(x\mid x_1)=\mathcal N(tx_1,(1-t)^2I)\), and the conditional velocity as a function of position is \(u_t(x\mid x_1)=(x_1-x)/(1-t)\). Plugging in a point on the path gives \(x_1-x_0\), matching the notes. This path is the \(\sigma_{\min}\to0\) case of Lipman et al.'s OT path and coincides with the linear interpolation of rectified flow (Liu et al., 2023).
The velocity field in Figure 2. When the data is a finite set \(\{x^{(i)}\}\), the average velocity can be computed exactly:
$$w_i\propto\exp\!\Big(-\frac{\lVert x-t\,x^{(i)}\rVert^2}{2(1-t)^2}\Big),\qquad u_t(x)=\frac{\sum_iw_i\,x^{(i)}-x}{1-t}.$$The figure is drawn from this formula, with no neural network. Following this optimum to the end always lands exactly on one of the training points. In other words, a model that drives the loss perfectly to its minimum only reproduces its training data, and real networks produce new samples precisely because they cannot follow this solution exactly. Also, as \(t\to1\) the denominator goes to zero and the field becomes singular.
Why trajectories never cross. If the velocity field is Lipschitz, ODE solutions are unique, so two particles at the same point at the same time must follow the same trajectory forever after. Trajectories from different starting points therefore cannot meet.
How is τ sampled during training?
π0 does not sample \(\tau\) uniformly between 0 and 1. It samples values near the noise end more often. For robot actions, the observation strongly narrows down the answer, so deciding the overall direction while there is still a lot of noise is harder than the final fine adjustments.
Go deeper The exact form of the distribution
With \(s=0.999\), \(p(\tau)=\mathrm{Beta}\big(\tfrac{s-\tau}{s};1.5,1\big)\). Normalized, \(p(\tau)=\tfrac{3}{2s}\sqrt{1-\tau/s}\ (0\le\tau\le s)\). Values \(\tau>s\) are discarded because, with an integration step larger than \(1-s\), that range is never evaluated at inference time.
5What makes it good
① A simpler prediction target
Because the path from noise to the answer is close to a straight line, the velocity field the model has to learn becomes much simpler and smoother.
Within a single pair, the training target \(\mathbf A-\epsilon\) is a constant that does not change over time. The problem handed to the model is simple, and that is a genuine advantage.
Another advantage is numerical stability. In fact, whether you predict the velocity, the clean action, or the noise, you can convert between them. But predicting the clean action or the noise and then converting to a velocity requires dividing by small numbers in some time ranges, which amplifies errors. Predicting the velocity directly gives the quantity needed for integration without any such division.
"The path is close to a straight line" connects to the catch in Section 4. The training target is straight, but the average velocity field the model learns is curved. Still, it is empirically known to be generally less curved than the paths diffusion produces, and that is the advantage we actually get.
Go deeper How the three prediction targets relate, and the score
At a fixed \((x_t,t)\), the velocity \(v\), the data estimate \(\hat x_1\) and the noise estimate \(\hat x_0\) are affinely related:
$$\hat x_1=x_t+(1-t)\,v,\qquad \hat x_0=x_t-t\,v.$$They read as "go on for the remaining time to reach the endpoint" and "go back for the elapsed time to reach the start." As losses, they differ only by a time-dependent weight:
$$\big\lVert v-(x_1-x_0)\big\rVert^2=\frac{\lVert\hat x_1-x_1\rVert^2}{(1-t)^2}=\frac{\lVert\hat x_0-x_0\rVert^2}{t^2}.$$Predicting \(\hat x_1\) and converting with \(v=(\hat x_1-x_t)/(1-t)\) amplifies errors as \(t\to1\); predicting \(\hat x_0\) and converting with \(v=(x_t-\hat x_0)/t\) amplifies them as \(t\to0\). This is the "division by small numbers" in the main text.
Relation to the score. Since \(\nabla_x\log p_t(x\mid x_1)=-(x-tx_1)/(1-t)^2=-x_0/(1-t)\), the score is \(s(x,t)=-\mathbb E[X_0\mid X_t=x]/(1-t)\), and
$$u_t(x)=\frac{x-\mathbb E[X_0\mid X_t=x]}{t}=\frac{x+(1-t)\,s(x,t)}{t}\quad(t>0).$$The velocity field is an affine function of the score, with coefficients that do not depend on the conditioning. This fact is used in ③ Guidance and in the comparison with diffusion in Section 6.
② Fast inference
Constant, smooth velocity means that with far fewer integration steps you get from noise to an accurate action trajectory almost instantly.
The direction is right, but there is a condition. Whether a few steps are enough depends on how straight the average trajectories are. If they are perfectly straight, one jump is exact; if they are curved, the fewer the steps, the more the samples miss the curve.
The extreme case is interesting. What happens with exactly one step? At the start there is no information about which answer a given noise sample will go to, so the velocity the model gives is "the average direction toward all answers." As a result, every sample collapses onto a single point, the mean of the data. Try increasing the number of steps from 1 below.
Euler trajectories using the optimal field
Final samples t = 1
Mean distance from each final sample to the nearest cluster center 0 (cluster standard deviation 0.18)
Translated to a robot, this is alarming. The average of a demonstration that "goes around the obstacle on the left" and one that "goes around on the right" is a trajectory that drives straight into the obstacle. π0's 10 steps are enough in most situations, but when several strategies are possible for the same observation, cutting the steps too aggressively risks producing this kind of averaged action.
"Reaching the target almost instantly" holds only when the trajectories are straight. This is why most research on reducing the step count focuses on straightening the trajectories.
Go deeper The size of the error and ways to straighten trajectories
The global error of Euler's method is \(O(\Delta t)\), with a constant proportional to the acceleration of the trajectory, \(\ddot x_t=\partial_tv+(\nabla_xv)v\). For straight trajectories \(\ddot x_t=0\), so a single step is exact.
The one-step collapse follows from this: at \(t=0\), \(X_0\) is independent of \(X_1\), so \(u_0(x)=\mathbb E[X_1]-x\) and therefore \(x_0+u_0(x_0)=\mathbb E[X_1]\).
Ways to straighten trajectories include reflow, which retrains on new (noise, output) pairs generated by the learned ODE (Liu et al., 2023); minibatch OT coupling, which pairs noise and data at minimum cost within a batch (Tong et al., 2024); and distilling a student model to produce the result of many steps in one.
③ Guidance
This part was left blank in my notes. Guidance is a family of techniques that slightly modify the velocity during sampling, after training is done, to strengthen some desired property. There are two kinds.
Following the condition more strongly. If the language instruction is occasionally dropped during training, the model learns both "the velocity with the instruction" and "the velocity without it." At sampling time, amplifying the difference between the two yields actions that follow the instruction more faithfully, at the cost of diversity.
$$\tilde v=v(\cdot,\varnothing)+w\,\big(v(\cdot,c)-v(\cdot,\varnothing)\big),\qquad w>1$$Enforcing constraints. Midway through integration, you can estimate "where we will land if we keep the current velocity to the end" (\(\hat{\mathbf A}=\mathbf A^\tau+(1-\tau)v\)). If this predicted endpoint collides with something or disagrees with actions already being executed, the velocity is corrected in the direction that reduces the problem. Real-time chunking by Black et al. (2025) uses this idea so that, while inference is running, the new chunk connects smoothly with the part of the previous chunk already being executed.
Constraint guidance does not guarantee that the constraint is satisfied. The predicted endpoint is not an actual sample but an average over possibilities; early in integration it is a blurry mean that can differ from where the sample actually ends up. For safety-critical constraints, you need separate verification or correction after sampling.
Go deeper Why this is exactly diffusion guidance
In the "Go deeper" box of ①, \(u_t=(x+(1-t)s)/t\) is an affine function of the score whose coefficients do not depend on the condition. A combination of velocities with weights summing to 1, \(\tilde v=(1-w)v_\varnothing+wv_c\), becomes \(\tilde v=(x+(1-t)\tilde s)/t\) with \(\tilde s=(1-w)s_\varnothing+ws_c\). This is the same operation as classifier-free guidance in diffusion (Ho and Salimans, 2022).
Constraint guidance for a cost \(J\) is written \(\tilde v(x_t,t)=v(x_t,t)-\lambda_t\nabla_{x_t}J\big(\hat x_1(x_t)\big)\) with \(\hat x_1=x_t+(1-t)v\). Since \(\hat x_1=\mathbb E[X_1\mid X_t=x_t]\), at small \(t\) it is a mixture-of-modes average, which is the precise reason for the limitation in the refinement box.
6Compared with diffusion
The first serious use of generative models for robot policies was Diffusion Policy (Chi et al., 2023). Diffusion repeats, dozens of times, "predict the mixed-in noise, subtract a little of it, then add a little fresh noise." Flow matching adds no fresh noise and moves deterministically along the velocity. Let us compare the two starting from the same noise and producing an action chunk.
Diffusion (DDPM) network calls 0 / 50
Flow matching network calls 0 / 10
The purple line (the path traced by one waypoint) makes the difference clear. Diffusion converges in a jagged way over 50 steps, while flow matching arrives in a straight line in 10. Keep in mind that this is the easy case of a single correct answer per observation, which is why the flow matching path is perfectly straight. With multiple correct answers it bends, as in Figure 2.
Go deeper They are different choices within the same framework
Linear interpolation is a special case of a more general Gaussian path:
$$x_t=\alpha_t x_1+\sigma_t x_0,\qquad u_t(x_t\mid x_0,x_1)=\dot\alpha_t x_1+\dot\sigma_t x_0.$$With \(\alpha_t=t,\sigma_t=1-t\) it is linear interpolation; plugging in the variance-preserving (VP) diffusion schedule gives the same probability path as diffusion (Lipman et al., 2023). DDPM's noise prediction (Ho et al., 2020) is \(\hat x_0\) prediction, which is score prediction. Song et al. (2021) showed that there is a deterministic ODE with the same distribution path as diffusion (the probability flow ODE), and DDIM is a way of integrating that ODE. Flow matching learns the velocity field of such an ODE directly by regression, without going through the score. In the end, the difference lies in the choice of path (\(\alpha_t,\sigma_t\)) and of prediction target.
7Summary
Connect noise and the answer with a straight line, learn the velocity along it by regression, and at inference time add that velocity a few times to produce an action.
| Stage | What happens | Key point |
|---|---|---|
| Step 1 | \(\mathbf A^0=\epsilon\sim\mathcal N(0,I)\), size \(H\times d_a\) | Not one action but a 50-step action chunk |
| Step 2 | Process images, language and state once with the VLM, cache the KV | The one-way mask makes caching exact |
| Step 3 | \(\mathbf A\leftarrow\mathbf A+0.1\,v_\theta\), ten times | The heavy VLM runs once; only the action expert repeats |
| Training | \(\lVert v_\theta-(\mathbf A-\epsilon)\rVert^2\) | The model learns the average velocity of the answers, so trajectories bend |
| Advantage ① | A simple, stable training target | The target is straight; the learned field is curved |
| Advantage ② | Few steps | Only when trajectories are straight; one step collapses to the mean |
| Advantage ③ | Guidance for conditions and constraints | Does not guarantee constraint satisfaction |
References
- Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, M. Le. Flow Matching for Generative Modeling. ICLR 2023. arXiv:2210.02747
- X. Liu, C. Gong, Q. Liu. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. ICLR 2023. arXiv:2209.03003
- M. S. Albergo, E. Vanden-Eijnden. Building Normalizing Flows with Stochastic Interpolants. ICLR 2023. arXiv:2209.15571
- A. Tong et al. Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport. TMLR 2024. arXiv:2302.00482
- J. Ho, A. Jain, P. Abbeel. Denoising Diffusion Probabilistic Models. NeurIPS 2020. arXiv:2006.11239
- Y. Song et al. Score-Based Generative Modeling through Stochastic Differential Equations. ICLR 2021. arXiv:2011.13456
- J. Song, C. Meng, S. Ermon. Denoising Diffusion Implicit Models. ICLR 2021. arXiv:2010.02502
- J. Ho, T. Salimans. Classifier-Free Diffusion Guidance. 2022. arXiv:2207.12598
- C. Chi et al. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. RSS 2023. arXiv:2303.04137
- K. Black et al. π0: A Vision-Language-Action Flow Model for General Robot Control. 2024. arXiv:2410.24164
- K. Black, M. Y. Galliker, S. Levine. Real-Time Execution of Action Chunking Flow Policies. 2025.