Source-linked AI summary

What Do Audio-Visual Synchronization Metrics Actually Measure?

Jai Kumar Sharma, Peeyush Tapadiya

arXiv:2608.25157v1cs.CVcs.MMcs.SD

TL;DR

Automatic AV-sync metrics are widely used to rank and train audio-visual generators, but their reliability as measurement instruments has not been jointly established. The paper audits four deployed metrics under controlled distortions, preprocessing tests, agreement analysis, PEAVS-proxy comparison, and fusion, finding an axis split rather than a single winner. It recommends reporting a Reliability Card instead of one bare synchronization score.

  • Problem

    Deployed AV-sync metrics are widely used for ranking and training, but they have rarely been jointly audited for reliability, stability, agreement, and human grounding.

  • Method

    The paper evaluates AV-Align, ImageBind AV-relevance, JavisScore, and Synchformer/DeSync with controlled desynchronization, preprocessing sensitivity, rank uncertainty, cross-metric agreement, PEAVS-proxy comparison, and learned fusion.

  • Results

    The audit finds an axis split: Synchformer leads temporal-offset tracking, ImageBind/JavisScore better match the PEAVS proxy and content-disruption families, and the metrics disagree with Krippendorff 𝛼=0.066.

  • Takeaways & Limitations

    AV-sync should be reported as a Reliability Card containing metric-family scores, uncertainty, and rank resolution rather than as a single bare number.

  • Takeaways & Limitations

    PEAVS is a learned proxy rather than fresh human labels, so the study treats this axis as PEAVS agreement rather than perceptual ground truth.

Abstract

from arXiv · show

Automatic AV-sync metrics are widely used to rank and train audio-visual generators, but they are rarely audited as measurement instruments. We jointly audit AV-Align, ImageBind AV-relevance, JavisScore, and Synchformer/DeSync under a common reliability protocol: controlled-distortion monotonicity, preprocessing sensitivity, rank uncertainty, cross-metric agreement, PEAVS-proxy agreement, and learned fusion. The result is an axis split, not a single winner: Synchformer/DeSync is the strongest temporal-offset tracker ($τ=0.84$), ImageBind/JavisScore better match the PEAVS human-aligned proxy ($τ=0.20$) and content-disruption families, and AV-Align is the weakest standalone metric. The metrics mutually disagree (Krippendorff $α=0.066$), and neither linear nor simple $k$-NN fusion improves PEAVS agreement over the best individual metric. We recommend reporting AV-sync as a Reliability Card (metric-family breakdowns with confidence intervals) rather than a single bare synchronization score.

1. Introduction

The paper argues that widely deployed AV-sync metrics should be audited as measurement instruments because unreliable scores can distort model ranking and optimization. It introduces a joint reliability audit spanning monotonicity, preprocessing stability, cross-metric agreement, perceptual alignment, and fusion.

  • AV-sync metrics are used to rank and train audio-visual generators, so unreliable scores can corrupt optimization.
  • The deployed metric family had not been jointly audited for stability, cross-metric agreement, rank uncertainty, and human grounding.
  • The audit applies a reproducible synthetic-desynchronization oracle, preprocessing-sensitivity harness, cross-metric and PEAVS-proxy comparisons, and learned fusion.
  • The headline result is an axis split: Synchformer tracks temporal offset best, embedding metrics align better with content disruption and PEAVS, and no metric wins both axes.

2. Related Work

Prior AV-sync work has largely validated individual metrics against their own design goals or used them for model ranking. This paper instead treats the deployed metrics as a group of black-box instruments requiring joint reliability evaluation.

  • AV-Align matches optical-flow motion peaks to audio-onset peaks using IoU, while Synchformer predicts audio-visual temporal offset.
  • ImageBind uses generic cross-modal cosine similarity, and JavisScore aggregates windowed ImageBind similarity for semantic correspondence.
  • ImageBind and JavisScore target semantic correspondence rather than onset-level timing, so weak temporal-oracle tracking is expected.
  • PEAVS trains a human-aligned predictor on 120K opinion scores, whereas prior studies generally validated one proposed metric or ranked models.

3. Method

The method evaluates each metric as a black box against controlled, known desynchronization and tests stability, inter-metric agreement, and human-aligned fusion. Kendall 𝜏 measures whether scores monotonically worsen as distortion increases.

  • The synthetic oracle applies increasing temporal shift, audio speed, fragment shuffle, and intermittent mute to real, well-synced clips.
  • Kendall 𝜏 compares metric scores with negative perturbation severity, where 𝜏=1 indicates perfectly monotonic degradation.
  • Preprocessing sensitivity is measured with coefficient of variation and rank-flip probability under resampling, cropping, and length truncation.
  • Cross-metric agreement uses pairwise Kendall 𝜏 and Krippendorff 𝛼 on per-metric z-scored clip scores, anchored by PEAVS human inter-annotator agreement.
  • A regularized linear model tests whether metric scores can be combined into a human-aligned PEAVS score.

4. Experiments

Experiments use AVSync15 clips and a reliability scorecard to compare metric behavior across controlled desynchronization families. The reported comparisons emphasize temporal tracking, content-disruption response, uncertainty, and preprocessing sensitivity.

  • AVSync15 contains 1,500 real highly-synced clips across 15 VGGSound classes, with the primary audit using 75 clips.
  • The scorecard reports higher Kendall 𝜏 as better, lower coefficient of variation and rank-flip probability as better, with confidence intervals in the supplement.
  • Figure 1 compares agreement with controlled desynchronization using Kendall 𝜏 and 95% confidence intervals.
  • Synchformer leads temporal and audio-speed tracking, while JavisScore and ImageBind respond more strongly to intermittent mute; AV-Align is weakest overall.

5. Results

The audit finds an axis split rather than a single winning metric: Synchformer leads temporal tracking, embedding metrics better capture content disruption, and AV-Align is weakest overall. Reliability also breaks down for close leaderboard gaps, while the metrics largely disagree about clip rankings.

  • Metrics split by reliability axis: 0.84 temporal-shift τ and 0.76 audio-speed τ make Synchformer the strongest tracker of pure temporal offsets.Its temporal-shift confidence interval is [0.78, 0.90], versus 0.16–0.39 for the other metrics.
  • Metrics split by reliability axis: 0.38 shuffle τ for ImageBind and 0.73 mute τ for JavisScore show that embedding metrics lead on content-disruption families.Synchformer scores 0.27 on shuffle and 0.59 on mute, so its temporal strength does not generalize across distortion types.
  • Preprocessing sensitivity and rank-flip: 0.54 median CV makes AV-Align the only high-variance metric, while every metric has a high worst-family adjacent flip probability.The reported maximum flip probabilities are 0.96, 0.69, 0.90, and 1.00 for AV-Align, ImageBind, JavisScore, and Synchformer.
  • Preprocessing sensitivity and rank-flip: 0.19 temporal-shift flip probability for Synchformer is lower than the others’ 0.46–0.66, but only Synchformer reaches d′ ≥1, with d′=1.6 on temporal shift.Its shuffle flip probability is 1.00, and leaderboard gaps below the 13–17% minimum detectable difference are noise.
  • Real generations: Only Synchformer reliably separates the close large-44k versus large-44k-v2 generators; similarity-metric rankings fall within noise.Far model gaps are robust across metrics, but close leaderboard gaps are not consistently resolvable.
  • Cross-metric disagreement: Krippendorff α=0.066 shows near-total disagreement among the four metrics, with ImageBind and JavisScore the sole strongly aligned pair at τ=+0.78.Synchformer agrees with the other deployed metrics at approximately τ=0, and the metrics span three orthogonal axes.

6. Discussion and Guidance

The audit shows that AV-sync metrics measure distinct axes rather than one shared construct: temporal tracking and PEAVS agreement are nearly orthogonal. The authors therefore recommend reporting a multi-axis Reliability Card instead of a single synchronization score.

  • The temporal and PEAVS-agreement axes are nearly orthogonal, with no metric performing highly on both.
  • PEAVS is a human-aligned learned proxy, not perceptual ground truth, and direct human preference annotation remains future work.
  • Krippendorff α=0.066 indicates that the deployed metrics mutually disagree, while the Reliability Card calls for reporting multiple axes rather than one bare number.

Appendix (supplementary material)

The supplementary analyses test whether the main findings depend on score reduction, implementation, setting, or preprocessing measurement. Across these checks, the temporal/content split and AV-Align’s weakness remain stable, motivating a reusable Reliability Card.

  • A. Synchformer scoring-reduction ablation: The temporal/perceptual split survives all three Synchformer score reductions: temporal-oracle τ stays high while PEAVS τ stays low.
  • B. AV-Align: official vs. our implementation: Official and reimplemented AV-Align show the same weak tracking, high variance, near-zero PEAVS agreement, and high temporal rank-flip.
  • C. Cross-metric agreement across settings: Agreement is near zero on clean clips but 3× higher on the controlled grid where synchronization varies by design.
  • D. Preprocessing-sensitivity measure: CV and rank-flip probability agree on preprocessing sensitivity, with AV-Align the outlier on both measures.
  • E. Metric Reliability Card: The Metric Reliability Card reports all seven axes instead of reducing AV-sync evaluation to a single number.
  • G. Second-domain replication: A VGGSound replication preserves the axis split: Synchformer leads temporal tracking, embedding metrics lead content disruption, and AV-Align remains weakest overall.
Loading 2608.25157v1…