Source-linked AI summary
When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?
Yansen Han, Hongxin Sun, Tao Lin
TL;DR
CFM losses are increasingly used as endpoint likelihoods and likelihood ratios in flow-matching alignment, but their exact validity is unclear. The paper derives a Gaussian-path decomposition into CFM and residual terms, showing when substitutions are exact and how biased ratios can still support optimization. Its conclusions separate off-policy population behavior from training and on-policy alignment, with numerical instability remaining in 32D.
Problem
The paper addresses whether conditional flow-matching losses and their old/new differences can validly replace pointwise endpoint NLLs and likelihood ratios needed for alignment.
Method
For linear Gaussian paths, the paper exactly decomposes endpoint NLL into entropy, weighted CFM, an interior velocity–score residual, and a boundary residual, then evaluates the theory empirically.
Results
At the fixed-target population optimum, ordinary CFM is not generally a pointwise NLL estimator, while wsc(t) = (1 − t)/t removes the interior residual; on-policy CFM-ratio surrogates can improve reward without being exact endpoint ratios.
Takeaways & Limitations
CFM substitutions should be treated as exact only when the relevant residuals cancel, and otherwise as controlled surrogates whose usefulness depends on optimization mechanisms.
Takeaways & Limitations
The analysis is limited to linear Gaussian paths with positive-time smoothing, and on-policy residual estimation remains numerically unstable in 32D.
Abstract
from arXiv · showhide
Flow matching enables likelihood-free training, yet alignment methods increasingly reuse conditional flow matching (CFM) losses as endpoint negative log-likelihoods (NLLs) and their old/new differences as log-likelihood ratios. We characterize when these substitutions are valid. For linear Gaussian paths, we exactly decompose endpoint NLL into entropy, a weighted CFM objective, an interior velocity--score residual, and a boundary residual. Thus CFM-only estimates and differences are exact only when the corresponding residuals cancel. At the off-policy population optimum, ordinary CFM is not generally a pointwise NLL estimator, whereas \(w_{\mathrm{sc}}(t)=(1-t)/t\) removes the interior residual; this positive result does not extend generally to training or on-policy alignment. On-policy log-ratios can remain biased even for identical endpoint laws or after surrogate optimization. Experiments across dimensions, distributions, and geometries support these conclusions and the mechanisms that make inexact ratios useful. **More broadly, the decomposition provides a theoretical basis for adapting likelihood-based LLM methods to flow matching, while distinguishing exact substitutions from controlled surrogates.**
1 Introduction
The paper asks when CFM-only quantities can replace pointwise endpoint likelihoods, especially in reward-based alignment where exact likelihood ratios are needed. It derives an exact Gaussian-path gap decomposition and separates conditions for valid substitutions from mechanisms making biased ratios useful.
- Flow matching replaces likelihood-based training with supervised regression, but alignment methods need per-sample likelihood ratios while endpoint density evaluation is expensive.
- Pointwise CFM averages velocity-regression errors along noisy paths, whereas endpoint NLL depends on the endpoint marginal, so population-optimal velocity need not imply pointwise equality.
- CFM-only estimates equal endpoint NLL when residuals cancel, while old/new CFM differences equal clean endpoint log-ratios when corresponding relative residuals cancel.
- For linear Gaussian paths, endpoint NLL decomposes into entropy, a weighted CFM objective, an interior velocity–score residual, and a boundary residual.
- The study combines decomposition verification across dimensions and distributions with off-policy, on-policy, and controlled empirical analyses of biased CFM-derived ratios.
2 Related Work
Related work connects likelihoods to continuous flows, score matching, diffusion objectives, and reverse-process ratios. This paper instead studies whether forward-process CFM surrogates recover clean endpoint likelihood ratios.
- Continuous normalizing flows and weighted score or diffusion objectives provide likelihood-related interpretations, but these do not make samplewise CFM equal to clean endpoint NLL.
- Reverse-process reinforcement learning uses discretized trajectory ratios that need not equal marginalized clean endpoint ratios.
- Forward-process methods construct CFM-based likelihood surrogates, including PPO-style ratios from differences of conditional flow-matching losses.
3 Preliminaries
The paper defines flow matching on a Gaussian interpolation between clean data and a standard normal base, then introduces weighted CFM objectives and clean endpoint ratios for alignment.
- The flow model uses a vector field vθ whose marginals satisfy the continuity equation, with p1 = N(0, Id) as base distribution and p0 as clean data.
- For the linear Gaussian path, Xt = (1 − t)X0 + tX1 with independent X0 and X1, the conditional law is Gaussian and the conditional velocity is defined analytically.
- The weighted conditional flow-matching objective compares the model velocity with the conditional velocity along the prescribed path.
- Ordinary CFM uses w(t) = 1, while score-calibrated CFM uses wsc(t) = (1 − t)/t.
- Alignment reweights rewards from the old model using a clean old/new endpoint likelihood ratio for importance weighting, clipping, or trust-region control.
4 A Pointwise Gap Decomposition Between NLL and CFM
This section establishes the decomposition framework for comparing pointwise NLL with CFM on a smoothed Gaussian conditional path. It identifies residual cancellation as the exactness criterion under stated regularity assumptions.
- The decomposition identity explains what must hold for pointwise CFM or CFM differences to represent clean endpoint NLLs or log-ratios.
- The analysis fixes a clean endpoint and uses a positive-time Gaussian path because the conditional distribution becomes a Dirac mass at t = 0.
- Velocity and score gaps between the conditional path and model are the quantities used to form the interior and boundary residual terms.
- Regularity, continuity, growth, and integration conditions support differentiation, integration by parts, and the endpoint limit required by the identity.
- The theorem states that pointwise endpoint NLL is a CFM term plus residuals, and the equation shows that the NLL is governed by a mixed velocity–score term rather than CFM alone.
5 Off-Policy Setting: Fixed-Target CFM Estimates
In the fixed-target setting, CFM estimates are exact only when interior and boundary residuals are negligible. At the population optimum, score calibration removes the interior residual, but a boundary term and training-time mismatch remain.
- During optimization: During fixed-target optimization, neither ordinary nor score-calibrated CFM is generally a training-time pointwise NLL estimator.The model is not constrained to satisfy the required pointwise velocity–score calibration relation during optimization.
- Exactness criterion: A CFM estimate is exact only when both interior and boundary residuals are negligible.This residual condition is the criterion separating an exact likelihood identity from an approximation.
- During optimization: A simple intermediate checkpoint demonstrates that even score calibration can fail during training.For a one-dimensional linear Gaussian path with a zero linear velocity coefficient, the pointwise calibration relation fails.
- After optimization: At the population optimum, wsc(t) = (1 − t)/t removes the interior velocity–score residual for any fixed endpoint distribution.The remaining discrepancy is the positive-time boundary term and finite numerical error.
6 On-Policy Setting: CFM-Ratio Surrogates
On-policy CFM-ratio surrogates replace clean endpoint log-ratios with differences of CFM objectives, but their exact bias is governed by relative residuals. These ratios can remain wrong despite unchanged endpoint laws or surrogate optimization, while clipping and reference control may support stable use.
- 6.1 The Exact Bias of a CFM-Only Ratio: Unclipped CFM-ratio objectives optimize a gap-tilted reward rather than necessarily the clean endpoint-ratio objective.The tilt is induced by the exact residual bias in the CFM-derived log-ratio.
- 6.1 The Exact Bias of a CFM-Only Ratio: A CFM-only ratio equals the clean endpoint ratio if and only if the relative residual ΔGε,w + ΔBε vanishes on update-relevant samples.The relative residual compares interior and boundary residual changes between new and old models.
- 6.2 Why the On-Policy Table Has No Global Identity Guarantee: Identical endpoint laws do not guarantee identical CFM ratios: the CFM log-ratio can be nonzero while the clean log-ratio is zero.This supplies a counterexample to any global on-policy likelihood-ratio identity.
- 6.2 Why the On-Policy Table Has No Global Identity Guarantee: Optimizing a regularized CFM-ratio surrogate need not eliminate the gap, since its optimizer can remain away from the zero-gap parameter.In the stated setting, if 0 < λ < rCε, the objective is maximized at a⋆ ≠ 0.
- 6.3 Why a Biased CFM Ratio May Still Be Useful: The decomposition establishes bias but not a stable optimization guarantee; clipping bounds proxy sample weights, while EMA and KL control reference distance without a proved residual bound.The experiments therefore test associations between these controls and optimization stability.
7 Experiments
Experiments test the decomposition identity, off-policy NLL replacement, on-policy ratio behavior, and stability mechanisms from 1D through 32D and across synthetic and image settings. They support the theory while exposing numerical and scope boundaries.
- RQ1: Across fixed-target synthetic evaluations, full decomposition estimates are more accurate than CFM-only endpoint estimates.In 1D, decomposition MAE ranges from 0.198–0.234 versus 0.222–0.803 for CFM-only estimates; in 2D, the ranges are 0.302–0.397 versus 0.342–1.587.
- RQ1: Decomposition identity error increases from about 1.02 in 4D to 2.62 in 32D.The remaining error reflects ODE, quadrature, score, and Monte Carlo approximation.
- RQ1: The 32D on-policy decomposition result is treated as failure of the available numerical estimator, not evidence against the analytic identity.Estimated residual terms reach the 10^9 scale in that setting.
- RQ2: wsc reduces fixed-target CFM-only NLL MAE from 0.800 to 0.223 in 1D, from 1.294 to 0.377 in 2D, and from 19.402 to 3.012 in 32D.The score-calibrated estimate is much closer when the boundary term is controlled, but two moons has a large finite-ε boundary term.
- RQ2: On-policy CFM-ratio training can achieve high reward alongside large clean-ratio errors.The displayed ordinary-CFM reward/source-relative-MAE pairs are 0.995/7.900 in 1D, 0.981/1.570 in 2D, and 0.910/1382.955 in 32D.
- RQ3: Reference updates are strongly associated with stability: full refresh is best in 1D/2D, moderate EMA is best in high-dimensional settings, and EMA-1 never passes the reward gate.Ratio clipping can help but tighter clipping is not uniformly better; these are controlled associations rather than causal evidence.
8 Conclusion
The paper establishes exactness criteria for replacing endpoint likelihoods with pointwise CFM quantities, separating off-policy population behavior from on-policy alignment. Experiments support score calibration for off-policy estimation while showing that useful on-policy surrogates need not be exact.
- The linear-Gaussian analysis decomposes endpoint NLL into entropy, weighted CFM, interior velocity–score, and boundary terms.
- CFM-only NLL estimates are exact when residuals vanish, while old/new CFM log-ratios are exact when corresponding relative residuals cancel.
- At the fixed-target population optimum, wsc(t) = (1 −t)/t removes the interior residual, but a boundary term remains; neither weighting has a general identity guarantee during optimization.
- On-policy CFM-ratio optimization can improve reward without ensuring an exact endpoint log-ratio, and successful optimization can coexist with large clean-ratio errors.
- The analysis is limited to linear Gaussian paths with positive-time smoothing, while 32D on-policy residual estimation remains numerically unstable.
A.3 Proof of Proposition 5.2
The appendix combines the population-optimal CFM characterization with numerical audits across dimensions and controlled on-policy experiments. These results support the decomposition, distinguish exact likelihood identities from useful surrogates, and expose numerical and methodological boundaries.
- Proof of Proposition 5.2: Every positive time weight has the same pointwise population CFM minimizer vθ⋆(x, t) = E[X1 −X0 | Xt = x].
- Proof of Proposition 5.2: The score-calibrated weight wsc(t) = (1 −t)/t removes the interior velocity–score residual at the population optimum, leaving the boundary term and numerical error.
- Numerical audits: Adding the independently estimated interior term produces a small smoothed-identity closure error, while remaining error reflects score, ODE, Monte Carlo, and quadrature approximations.
- On-policy results: On-policy CFM-only ratios can achieve high reward, but nonzero source-relative MAE shows they are not exact on evaluated samples.
B.4.4 RQ3: Why do on-policy CFM-ratio surrogates work?
The controlled sweeps show that on-policy CFM-ratio optimization can remain useful despite inexact ratios, with stability depending strongly on reference updates and clipping.
- Stability criterion: CFM-ratio training is considered stable when nonzero ratio error does not prevent reward improvement.The sweeps measure control–stability associations, not a residual-versus-distance theorem.
- Other mechanisms: Changing ratio weights, advantage clipping, source-KL, or EMA-KL did not independently rescue experiments with EMA = 0.5 or EMA = 1.Each mechanism succeeded broadly at EMA = 0 but failed to repair the larger tested EMA settings by itself.
- Ratio clipping: The ratio-clip sweep produced 5/8 passing settings at EMA = 0.5 and 2/8 at EMA = 1, with non-monotonic effects across widths and variants.It was the only tested mechanism sweep with successful settings beyond EMA = 0.
- Reference updates: EMA = 0 is most consistently associated with passing the 0.90 reward gate, while ratio clipping provides the most passing configurations beyond EMA = 0.This is an empirical association from one seed, not a causal or local-accuracy result.
- Experimental design: The 2D audits vary EMA distance, estimate weight, ratio clipping, advantage clipping, source KL, and EMA KL using off-policy checkpoints with train weight w = 1.The runs use a 0.90 reward gate and report controlled settings separately.
B.5.3 RQ2: Do the table conclusions hold beyond 1D?
The 2D results support the paper’s off-policy and on-policy conclusions beyond 1D: score calibration reduces some endpoint errors, while CFM-only ratios remain inexact despite reward gains.
- Off-policy: The off-policy conclusions in Tab. 1 hold for 2D settings.This extends the reported off-policy distinction beyond the 1D experiments.
- Off-policy: During off-policy optimization, all three target families retain non-negligible endpoint-estimation error under both training and readout weights.The metric is MAE(NLL, Jw + H).
- Off-policy: After optimization, score-calibrated readout reduces direct endpoint MAE from 1.001 to 0.411 on GMM and from 1.587 to 0.342 on banana.The two-moons result shows score calibration alone is insufficient when the finite-ε boundary term is large.
- On-policy: The on-policy conclusions in Tab. 1 hold for 2D settings, including high reward from CFM-only ratio surrogates.The results are reported for the tested 2D configurations.
- On-policy: Nonzero source-relative MAEs show that 2D CFM-only ratios are not exact on evaluated samples, although round-old MAEs are smaller in these runs.The smaller round-old errors are an association, not evidence that reference closeness forces residual cancellation.
B.5.4 RQ3: Which controls are associated with stable 2D CFM-ratio updates?
In the tested 2D CFM-ratio sweeps, small or zero EMA decay was most consistently associated with passing the reward gate, while ratio clipping helped some larger-EMA settings non-monotonically. These single-seed associations do not establish causality or clean-ratio calibration.
- EMA distance: EMA = 0 was most consistently associated with passing the 0.90 reward gate across the tested 2D configurations.The reported fractions vary by sweep, including 4/6 successes for EMA = 0 and no successes at EMA = 0.5, 0.9, or 1 in one sweep.
- Ratio clipping: The ratio-clip sweep was the only controlled sweep with passing EMA > 0 configurations, but its effect was non-monotonic across clipping widths and variants.At EMA = 0.9, δ = 0.2 and no clipping pass 5/6 and 6/6 settings, respectively; at EMA = 1, all four successes occur with no clipping.
- Advantage clipping: Advantage clipping produced 20/24 successes at EMA = 0 but no successful setting at EMA = 0.5, 0.9, or 1.This supports an association with low EMA decay rather than a general benefit across reference-update distances.
- KL regularization: Positive source KL penalties eliminated three non-finite EMA-0 outcomes and yielded 14/18 successes, but no large-EMA experiment succeeded.EMA-KL regularization yielded 19/24 successes at EMA = 0, with none at EMA = 0.5, 0.9, or 1.
- Interpretation: The reported single-seed configuration fractions establish neither a causal stabilization mechanism nor a local or global clean-ratio calibration guarantee.The 2D sweep tables aggregate best rewards and passing/tested configurations rather than repeated-seed success probabilities.
- Scaling: Across dimensions, score-calibrated weighting became increasingly fragile, while high-dimensional training was more stable with EMA = 0.5 or 0.9 than with EMA = 0 or 1.The paper attributes this fragility to the score-calibrated weight enlarging the noise scale and attributes instability to the training objective rather than ratio estimation.