Source-linked AI summary
SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation
Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park
TL;DR
Whole-frame metrics can conflate desirable natural motion with background drift or flow stagnation in long-horizon fixed-camera scenes. SNF-Bench separates static support from dynamic flow, validates its factors through controlled corruptions, and audits released generators. The audit shows that spatially resolved measurements can reverse the interpretation and ordering produced by conventional whole-frame metrics.
Problem
Whole-frame metrics do not distinguish desirable motion in dynamic regions from erroneous motion in static background regions or from flow decay.
Method
SNF-Bench partitions scenes into static support and dynamic flow, reports static fidelity and flow persistence separately with drift leakage as context, and validates factor selectivity using known corruptions.
Results
The public-model audit shows that whole-frame metrics and SNF-Bench can produce materially different interpretations and orderings of the same long-horizon outputs.
Takeaways & Limitations
SNF-Bench makes the spatial origin and temporal persistence of motion explicit without collapsing them into a composite score.
Takeaways & Limitations
SNF-Bench is limited to fixed-camera generation with separable static support and dynamic flow, and it does not measure physical or semantic realism or distinguish progression from repetition.
Abstract
from arXiv · showhide
Long-horizon video generation is evaluated with whole-frame metrics that reward motion and temporal consistency. For fixed-camera nature scenes this creates an ambiguity: motion of water, fire, smoke, or rain is desirable, whereas motion of the background is an error. A system can therefore score well on motion while its scene drifts, or on consistency while its flow stagnates. We introduce SNF-Bench, an evaluation framework for long-horizon fixed-camera generation that partitions each scene into static support and dynamic flow and reports static fidelity, flow persistence with absolute magnitude, and drift leakage separately, never as one score. Drift leakage is interpretive context rather than a headline measurement. Each factor is validated mechanistically rather than by correlation with preference: we inject global translation, rotation, and scale drift and progressive late freezing at known severity into real generations, and require each factor to respond in its stated direction and to remain selective against corruptions it does not target. Auditing publicly released long-horizon text-conditioned checkpoints under one recorded common inference configuration, plus an image-conditioned track with released-pipeline references and a deployment-sensitivity panel, we find that whole-frame motion and static-region drift induce near-opposite orderings of the same outputs. At maximum controlled translation, fBD and NBF rise to $1.86\times$ and $1.32\times$ baseline, but whole-frame Dynamic Degree reaches only $1.07\times$---rewarding the corruption. SNF-Bench measures where motion occurs and whether it persists; it does not measure physical realism. Project page: https://minar09.github.io/snfbench/.
1. Introduction
SNF-Bench addresses the ambiguity of whole-frame evaluation in long-horizon fixed-camera nature videos by separating static support from dynamic flow. It reports static fidelity, flow persistence, and contextual drift leakage separately, with mechanistic validation and a public-model audit showing that spatially resolved measurements can change output interpretation.
- Long-horizon errors accumulate into background displacement, color drift, and motion decay that whole-frame metrics compress into one score.
- Fixed-camera nature scenes require stationary banks, rocks, and landscapes to remain fixed while water, fire, smoke, rain, or vegetation motion persists.
- SNF-Bench evaluates 60 and 120 s horizons using static fidelity and flow persistence as primary axes, while treating drift leakage as interpretive context.
- DD 0.908 vs 0.919 receives nearly identical whole-frame scores even though the corresponding outputs differ in background drift by over 40×.
- Controlled translation, rotation, scale drift, motion attenuation, freezing, photometric drift, and repetition test whether each factor responds selectively to its targeted corruption.
- The benchmark’s public-model audit finds high whole-frame motion can coexist with static displacement, while high stability can coexist with decaying flow.
3. SNF-Bench Design
SNF-Bench evaluates fixed-camera long-horizon video by separating stationary scene support from persistent natural flow. It reports static fidelity and flow persistence as distinct axes, uses drift leakage only as context, and validates measurements against controlled corruptions.
- Design: SNF-Bench treats fixed-camera videos as compositions of stationary support and dynamic flow, isolating static drift from flow decay.The framework evaluates both text- and image-conditioned generation under the same decomposition, with static support established from the opening sequence or a shared source image.
- Evaluation scope: Conclusions rest on the 60 s horizon, with 120 s as a supporting tier and 5 s and 240 s as diagnostic regimes.Runs that cannot reach a target horizon under their declared setting are reported as unsupported rather than omitted.
- Region partition: The automatic partition thresholds early optical flow into dynamic and static regions, then erodes the static mask and discards a 4% border as an ignored transition band.The partition is recomputed under systematic erosion and dilation because early drift can influence the measured regions.
- Design: The benchmark primarily reports Static Fidelity and Flow Persistence, while Drift Leakage contextualizes them rather than forming a third measurement.It avoids a single scalar because freezing the entire video can reduce background motion, while whole-scene displacement can create high foreground motion.
- Static fidelity: Feature-aligned Background Drift measures late geometric displacement in nominally static structure, while Normalized Background Flow measures static-mask flow normalized for resolution and frame rate.fBD uses matched ORB features; NBF averages optical-flow magnitude inside the static mask.
- Flow measurement: Global similarity-transform compensation removes coherent translation, rotation, and scale before measuring Motion-Compensated Foreground Flow.A translation-only estimate would leave rotational and zoom corruption in the dynamic-region measurement; median translation is used only as a recorded fallback when correspondences are insufficient.
- Flow persistence: Flow Persistence is the clipped late-to-early ratio of compensated foreground flow and is always reported with absolute MCFF magnitudes.MCFF-E and MCFF-L prevent uniformly negligible motion from appearing strong through a relative measure alone.
- Drift leakage: Drift Leakage Ratio compares static-region flow with raw dynamic-region flow but is deliberately unbounded and is not a causal decomposition of motion.High DLR with moderate DAR indicates static-region motion unexplained by the rigid model, without identifying the residual’s source.
5. Controlled Validation
SNF-Bench is validated with controlled corruptions to test whether each factor responds monotonically, selectively, and in its defined direction. Translation increases static-drift measures, while late freezing sharply reduces flow-persistence measures; partition checks preserve the headline NBF ordering but reveal an open area–fBD validity concern.
- Perturbations and Criteria: Controlled variants isolate geometric drift, motion attenuation, and late freezing at known corruption types and severities.Geometric drift applies increasing translation, rotation, or scale; attenuation and freezing weaken or remove later dynamic content.
- Perturbations and Criteria: Admission requires each factor to respond in the expected direction, monotonically with severity, selectively against unrelated corruptions, and robustly to partition changes.Table 1 sets admission at ρ ≥0.7 with off-target response below 25%.
- Validation Results: At maximum translation, fBD and NBF reach 1.86× and 1.32× baseline, while VBench Dynamic Degree reaches only 1.07×.The whole-frame metric therefore changes little and slightly upward under a corruption that fixed-camera evaluation should penalize.
- Validation Results: Late freezing reduces MCFF-L and FP to 0.02 and 0.07 of unperturbed values while leaving static factors comparatively unchanged.This selective response supports separating static fidelity from flow persistence.
- Validation Results: Rotation weakens both static-fidelity responses, while DLR reverses under rotation and is retained only as a comparative diagnostic alongside DAR.The reported rotational changes are +0.70 for fBD, +0.60 for NBF, versus +1.00 for both under translation; looping content remains a scope boundary because it can sustain persistence and magnitude through repetition.
- Validation Results: The headline NBF ordering is unchanged when precipitation is excluded, but the negative area–fBD association remains an open validity concern.Its direction is consistent with early-mask circularity, so the concern is reported rather than dismissed.
6. Public-Model Audit
The public-model audit compares released long-horizon generators in a shared recorded T2V configuration using paired-item evaluation and whole-frame baselines. Its analysis is designed to test whether SNF-Bench changes interpretation of the same outputs without treating stochastic trajectories as equivalent across checkpoints.
- Public-Model Audit: The audit asks how checkpoints occupy the static-fidelity/flow-persistence space rather than constructing a universal leaderboard.Seven publicly released autoregressive or streaming generators with accessible checkpoints were frozen before the sweep.
- Public-Model Audit: Every T2V checkpoint is evaluated through one recorded inference implementation with a shared four-step denoising schedule, six frames per block, and seed 0.The configuration uses scheduler-warped indices and a common step budget for the distilled checkpoint family.
- Public-Model Audit: Identical conditioning and fixed seeds make runs reproducible, but paired-item comparisons do not make different checkpoints’ stochastic trajectories equivalent.Scores are computed per item and averaged, with headline separation claims based on paired bootstrap distributions of per-prompt differences.
- Public-Model Audit: The same outputs are additionally scored with established whole-frame metrics, and the resulting orderings are compared with SNF-Bench interpretations.This directly tests whether the decomposition changes how model outputs are read.
7. Results and Interpretation Changes
The audit separates static fidelity and motion persistence, revealing distinct operating regimes and deployment-sensitive comparisons rather than a single ranking.
- Recorded Common-Configuration T2V Audit: Sustained flow requires non-trivial late motion and high retention; MCFF-E beside MCFF-L exposes uniformly negligible motion.
- Recorded Common-Configuration T2V Audit: Infinite-Forcing has the lowest NBF but the least surviving late motion, while Causal-Forcing occupies the higher-background-flow corner.Paired per-prompt intervals versus all six peers exclude zero for Infinite-Forcing’s NBF under the recorded configuration.
- Recorded Common-Configuration T2V Audit: The primary T2V audit reports static fidelity and flow persistence separately, with MCFF-E→L accompanying FP and DLR/DAR supplementary.Table 3 uses a recorded common configuration; 60 s is primary and 120 s is a directional check.
- Whole-Frame Motion and Static-Region Flow Yield Different Interpretations: Dynamic Degree ranks Causal-Forcing 1/7 and Infinite-Forcing 7/7, while NBF reverses those rankings across seven checkpoint means.The descriptive correlation between Dynamic Degree and static-region flow is ρ = +0.93.
- Whole-Frame Motion and Static-Region Flow Yield Different Interpretations: Figure 3 orders columns by increasing fBD across t=0,30,60 s for one shared prompt, highlighting the CausVid/Causal-Forcing pair.
- Image-Conditioned Track: I2V results keep released pipelines separate from wrapper outputs at 60 and 120 s, providing deployment sensitivity without cross-setting ranking.Figure 4 uses a distinct 60 s sample with a shared source image.
8. Limitations and Conclusion
SNF-Bench is scoped to fixed-camera scenes where static support and dynamic flow separate, and it concludes that region-resolved reporting complements broad benchmarks without a composite score.
- Limitations: SNF-Bench is limited to long-horizon fixed-camera generation where static support and dynamic flow separate.
- Limitations: The benchmark measures neither physical nor semantic realism, cannot distinguish progression from repetition, and depends on an automatic partition bounded by erosion–dilation analysis.
- Conclusion: The conclusion recommends separate static-fidelity and flow-persistence reporting, with drift quantities treated as context rather than collapsed into a composite score.Independent partitions and broader non-water coverage remain identified as next steps.
- Formal Factor Definitions: Feature-aligned Background Drift uses matched ORB pairs and expresses displacement as a percentage of image size.
- Formal Factor Definitions: Normalized Background Flow uses sampled flow intervals normalized by frame width, making it independent of native frame rate and sampling policy.
- Formal Factor Definitions: Global-motion compensation fits a robust similarity transform to static-region correspondences and subtracts its induced displacement field from optical flow.
- Formal Factor Definitions: Drift leakage and attenuation use static, raw dynamic, and compensated late-window flow magnitudes; DLR is unbounded, while DAR is clipped for reporting.DLR > 1 means static support moves more than the intended dynamic region.
S2. Region Partition
SNF-Bench derives a static-support/dynamic-flow partition from early-window optical flow, excludes uncertain boundaries, and validates its factors against controlled corruptions. The approach reduces circularity from late drift but remains subject to early-window and annotation limitations.
- Region Partition: Mean optical-flow magnitude over the leading 12% of sampled frame pairs is thresholded by Otsu’s method into dynamic and complementary static regions.The static region is eroded and a 4% frame border is discarded, creating an ignored transition band.
- Region Partition: The partition is fixed from an early window because using the full motion field would let accumulated background drift define the measured dynamic region.At the 60 s horizon, the partition is fixed from the first 7.1 s.
- Validation: Admission requires signed target-family Spearman correlation ρ ≥0.7 and off-target response below 25%.Rotation is reported but not gated, with DLR inverted and fBD and DAR carrying that case.
- Limitations: The automatically derived partition has no inter-annotator agreement, and early drift can influence the partition itself.The paper recomputes every measurement under systematic erosion and dilation of the boundary.
S3. Validation Protocol
The validation protocol tests whether partition-based measurements respond to intended corruptions while remaining comparable across geometric perturbation families. Robustness checks examine mask advantage, support-area comparability, and the two-way partition’s inability to represent crossing precipitation.
- Validation Protocol: Geometric corruption severity is normalized by mean pixel displacement over the frame at rollout end.Translation uses d pixels, rotation uses d/¯r radians, and zoom uses 1 + d/¯r, where ¯r is mean distance from frame center.
- Validation Protocol: The protocol asks whether systems can benefit from easier masks and whether rankings depend on precipitation, the category a two-way partition cannot represent.This frames partition validity as a robustness question rather than assuming the segmentation is neutral.
- Partition Robustness: The static region occupies 0.61–0.66 of the frame on average across audited systems, so support size is not markedly different between systems.Across seven systems, support area correlates negatively with feature-based drift, but the correlation is treated as an unresolved circularity risk.
- Composition Robustness: Removing precipitation leaves normalized-background-flow ordering unchanged and shifts only midfield positions under feature-based drift.This tests whether the unrepresentable crossing-flow category determines the main ordering.
S5. Paired Differences
Paired per-prompt differences, rather than marginal interval overlap, provide the basis for separation claims under a recorded common configuration. The audit compares checkpoints and tracks deployment sensitivity without treating the common schedule as each method’s own operating point.
- Paired Differences: Separation claims use bootstrap distributions of paired per-prompt differences because all systems receive the same prompts.Marginal interval overlap can conceal a reliable paired difference, while non-overlap can suggest a difference that is not present.
- Generation Configuration: All text-conditioned checkpoints use the same recorded four-step denoising configuration, six frames per block, seed 0, and 832 × 480 output at 16 fps.Within a horizon, checkpoint weights and recorded regular/EMA selection vary.
- Scope: The common-configuration audit is reproducible but does not measure methods at their released defaults or necessarily at their own operating points.Re-running each method at released defaults is identified as the natural next audit version.
- Image-Conditioned Track: Image-conditioned dagger rows use a common rollout wrapper and are interpreted as deployment sensitivity rather than ranked against released-pipeline outputs.Their denoising schedules are checkpoint-specific.
- Paired Differences: −83.98 is the Infinite-Forcing −Causal-Forcing mean difference in the 60 s NBF comparison, with a 95% CI of [−102.59, −65.52].The comparison uses n = 23 paired items.
S9. Aggregation Sensitivity
Aggregation sensitivity compares prompt-weighted and category-balanced views while preserving the reported roster and track distinctions. The best and worst systems remain unchanged, but intermediate ordering is not treated as stable.
- Aggregation: Prompt averaging is preferred because category balancing would give two-prompt cells the same influence as a fourteen-prompt cell and increase variance.The evaluation set contains five populated categories, including three with two prompts each.
- Aggregation: The best- and worst-ranked systems are identical under both aggregation schemes, while intermediate ranks move.Consequently, the paper makes no ordinal claim about the middle of the field.
- Roster: Causal-Forcing++ and other listed systems span released and wrapper settings across 5 s, 60 s, 120 s, and 240 s horizons.The roster includes both 1-step and 2-step Causal-Forcing++ variants.
- Measurement Coverage: fBD is withheld when too few repeatable keypoints survive in the static region, with public-roster abstentions confined to a desert dust-storm scene.Other factors remain reported for those clips.
S11. Deployment Sensitivity
The deployment-sensitivity analysis compares released-pipeline outputs with a common long-horizon wrapper and reports changes without treating wrapper scores as published performance. It also documents signed negative DAR values and scene-category coverage.
- Table S13 reports changes from released-pipeline outputs to the common long-horizon wrapper for checkpoints evaluated under both settings.
- The analysis reports the incidence and magnitude of negative signed DAR values rather than suppressing them.
- Negative signed DAR values are informative because opposing local flow can increase measured magnitude after global-motion compensation.
- Scene-category coverage is summarized by track and horizon.
S14. Additional Qualitative Material
Supplementary material documents audit coverage, comparison layouts, rank disagreement, deployment sensitivity, and scope boundaries for SNF-Bench. It also clarifies how selected prompts and factor profiles should be interpreted across text- and image-conditioned tracks.
- Audit coverage: The supplementary audit tables cover complete T2V and I2V results across four horizons, with frozen specifications and explicit limits on cross-horizon and cross-setting comparisons.The T2V and I2V tables distinguish initialization, principal, and diagnostic horizons; I2V released-pipeline and wrapper rows are reported separately.
- Track interpretation: Image-conditioned comparisons use one shared source image, whereas text-conditioned systems author their own layouts; released-pipeline and wrapper outputs are therefore separated.The figure explicitly warns against cross-setting ordering and distinguishes wrapper outputs with †.
- Factor profiles: Factor profiles orient every factor so outward is better, but no system encloses the others on either track, supporting an operating-regime rather than single-ranking view.Raw values would make high drift and surviving motion point outward simultaneously, so the profile uses rescaled, direction-adjusted factors.
- Scope boundaries: The stated scope boundaries include precipitation-support entanglement, stagnation that can mimic static fidelity, and residual motion whose composition is not identified.High leakage with low attenuation indicates static-region motion, but these measurements do not determine what the residual contains.
- Rank comparisons: Rank disagreement is assessed by comparing whole-frame motion with normalized background flow, while the 60 s table reports descriptive Spearman correlations over n = 7 method means.The supplementary material frames positive correlation between apparent motion and background flow as the effect isolated by the benchmark.