Flow matching models learn the time-dependent vector field that transports samples from a known distribution
Let’s use scalar interpolation schedules
We assume that
Conditioning on the endpoint
The pathwise velocity is
For interior times at which
The corresponding conditional vector field is therefore
Assuming
The conditional training objectives recover the corresponding marginal fields through posterior averaging:
Both posterior averages can be expressed using the posterior mean of the data endpoint, which we call the denoiser:
Taking conditional expectations of the affine path and its velocity gives
Eliminating the posterior mean of
Similarly, posterior averaging of the conditional Gaussian score gives the marginal score
and therefore
Thus, at interior times where
The velocity can be learned with the conditional flow matching loss, while the score can be learned with the conditional score matching loss. The denoiser can likewise be learned by regressing directly onto
Now that we have a score function, for any choice of diffusion coefficient
To see why this SDE preserves the same probability path as the original flow ODE, consider its probability current:
Here we used
When initialized from the same
These noise-to-data dynamics can be viewed as the reverse-time dynamics of a corresponding forward noising process, allowing us to construct stochastic paths transporting
The animations below compare the deterministic flow ODE with isotropic, anisotropic, and state-dependent SDEs that share its marginal probability path. The arrows can display the ODE velocity