Source-linked AI summary

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

Kai Liu, Wei Li, Lai Chen, Shengqiong Wu, Yanhao Zheng, Jiayi Ji, Fan Zhou, Jiebo Luo, Ziwei Liu, Hao Fei, Tat-Seng Chua

arXiv:2503.23377v2cs.CVcs.AIcs.SDeess.AS

TL;DR

JAVG requires high-quality audio and video generation with precise synchronization, but existing methods and benchmarks are limited for fine-grained alignment and complex real-world scenes. JavisDiT addresses this with hierarchical spatio-temporal priors, while JavisBench and JavisScore support broader evaluation. Experiments report state-of-the-art generation and synchronization performance, with video quality remaining slightly below UniVerse-1 because of the weaker OpenSora backbone.

  • Problem

    Existing JAVG methods have limited fine-grained modeling or cross-modal information exchange, while common benchmarks contain simplistic content and limited scene diversity.

  • Method

    JavisDiT jointly generates audio and video with DiT branches guided by global and fine-grained spatio-temporal priors from HiST-Sypo, alongside JavisBench and JavisScore.

  • Results

    JavisDiT achieves state-of-the-art performance across existing closed-set and open-world benchmarks for content generation and audio-video synchronization.

  • Takeaways & Limitations

    JavisDiT establishes a benchmarked approach for synchronized audio-video generation in diverse and complex real-world scenarios.

  • Takeaways & Limitations

    Video quality is slightly lower than UniVerse-1, mainly because JavisDiT relies on the weaker OpenSora-1.2 backbone compared with Wan-2.1.

Abstract

from arXiv · show

This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates high-quality audio and video content from open-ended user prompts in a unified framework. To ensure audio-video synchronization, we introduce a fine-grained spatio-temporal alignment mechanism through a Hierarchical Spatial-Temporal Synchronized Prior (HiST-Sypo) Estimator. This module extracts both global and fine-grained spatio-temporal priors, guiding the synchronization between the visual and auditory components. Furthermore, we propose a new benchmark, JavisBench, which consists of 10,140 high-quality text-captioned sounding videos and focuses on synchronization evaluation in diverse and complex real-world scenarios. Further, we specifically devise a robust metric for measuring the synchrony between generated audio-video pairs in real-world content. Experimental results demonstrate that JavisDiT significantly outperforms existing methods by ensuring both high-quality generation and precise synchronization, setting a new standard for JAVG tasks. Our code, model, and data are available at https://javisverse.github.io/JavisDiT-page/.

1 INTRODUCTION

JAVG addresses the need to generate audio and video jointly because they are interconnected in real-world scenarios, while existing approaches face limitations in fine-grained modeling and mutual information exchange. JavisDiT combines DiT-based generation with hierarchical spatio-temporal priors, introduces JavisBench and JavisScore, and reports strong performance on existing and open-world benchmarks.

  • Audio-video joint generation is valuable because audio and video are inherently interconnected in most real-world scenarios.
  • Existing DiT-based JAVG methods can limit fine-grained spatio-temporal modeling or lack sufficient mutual information exchange between audio and video channels.
  • JavisDiT uses shared DiT-based audio and video branches plus HiST-Sypo to extract global and fine-grained spatio-temporal priors for synchronization.The priors capture overall semantic structure, sound-triggered visual content, and corresponding temporal information.
  • JavisBench contains 10,140 high-quality text-captioned sounding videos spanning five dimensions and 19 scene categories, with over half featuring complex scenarios.JavisScore is also introduced to evaluate synchronization on complex sounding videos.
  • JavisDiT significantly outperforms state-of-the-art methods across existing benchmarks and JavisBench, especially on complex-scene videos.The authors attribute its complex-scene effectiveness to the HiST-Sypo estimation mechanism.

2 PRELIMINARIES

JAVG models jointly denoise video and audio from a text input. The target is both coarse semantic alignment with the text and fine-grained spatial and temporal alignment between modalities.

  • JAVG diffusion models generate a video v and corresponding audio a simultaneously from text input s.
  • The diffusion update maps noisy video and audio states at timestamp t to preceding states through Gθ(vt, at, s, t).
  • An ideal JAVG model preserves coarse semantic alignment with the input text while achieving fine-grained spatiotemporal alignment between video and audio.
  • Spatial alignment matches events at video-frame regions with corresponding frequency components in the audio spectrogram.
  • Temporal alignment requires event onset or termination in video to coincide with the corresponding audio response start or stop.

3 THE PROPOSED JA V I SDIT SYSTEM

JavisDiT uses a DiT-based dual-branch architecture with spatio-temporal attention, hierarchical priors, and bidirectional audio-video interaction. Its prior estimator learns variable event locations and timings, while staged training supports joint synchronized generation.

  • 3.1 JAVISDIT MODEL ARCHITECTURE: Cascaded spatial-then-temporal self-attention aggregates intra-modal information while reducing computational cost.
  • 3.1 JAVISDIT MODEL ARCHITECTURE: JavisDiT’s architecture contains video and audio generation branches, HiST-Sypo, and multimodal bidirectional cross-attention.The detailed blocks include ST-SelfAttn, Fine-Grained ST-CrossAttn, and MM-BiCrossAttn.
  • 3.1 JAVISDIT MODEL ARCHITECTURE: Spatial and temporal priors guide cross-attention in both branches, providing unified fine-grained conditioning for audio-video synchronization.
  • 3.1 JAVISDIT MODEL ARCHITECTURE: Bidirectional cross-attention enables direct audio-video interactions by exchanging information in both directions.
  • 3.2 HIERARCHICAL SPATIAL-TEMPORAL SYNCHRONIZED PRIOR ESTIMATOR: HiST-Sypo combines a global semantic prior for what occurs with a fine-grained spatio-temporal prior for when and where it occurs.
  • 3.2 HIERARCHICAL SPATIAL-TEMPORAL SYNCHRONIZED PRIOR ESTIMATOR: The prior estimator uses ImageBind text features and learnable spatial and temporal tokens to sample plausible priors from a Gaussian distribution.Contrastive learning with asynchronous negative pairs trains the estimator to capture variability in event timing and location.
  • 3.3 MULTI-STAGE TRAINING STRATEGY: JavisDiT follows audio pretraining, ST-Prior training, and JAVG training stages to support single-modal quality and synchronized joint generation.
  • 3.3 MULTI-STAGE TRAINING STRATEGY: Dynamic temporal masking adapts JavisDiT to video-to-audio, audio-to-video, image animation, and video-audio extension tasks.

4 A CHALLENGING JA V I SBENCH BENCHMARK

JavisBench is designed to address the limited diversity and realism of existing JAVG benchmarks through a hierarchical taxonomy, curated sounding videos, and evaluation of complex synchronization scenarios. It also introduces JavisScore to measure spatiotemporal audio-video synchronization in real-world content.

  • Motivation: Existing JAVG benchmarks lack diverse scenes, complex real-world environments, and multiple sounding events within individual audio-video samples.These limitations constrain comprehensive and realistic evaluation of joint audio-video generation.
  • Taxonomy: JavisBench organizes evaluation across five dimensions: event scenario, video style, sound type, spatial composition, and temporal composition.The dimensions range from coarse scenario descriptions to fine-grained spatial and temporal relationships between sounding events.
  • Taxonomy: GPT-4 supports a hierarchical categorization system containing 5 dimensions and 19 categories for JavisBench.The taxonomy is used to structure the benchmark’s detailed category distribution.
  • Data Construction: JavisBench combines existing dataset test sets with YouTube videos uploaded between June and December 2024, collected using taxonomy-specific keywords and manually verified.The collection process is intended to target relevant categories while addressing data leakage and noisy data concerns.
  • Benchmark Statistics: JavisBench provides more diverse data than AIST++ and Landscape, a more detailed taxonomy than TAVGBench, and the first benchmark focused on multi-event synchronization.Its distribution covers event scenarios, visual styles, and audio types, including under-represented realistic scenarios.

5 EXPERIMENTS

Experiments show that JavisDiT achieves strong generation quality and audio-video synchrony across open-world and closed-set benchmarks, while remaining challenged by complex multi-object and multi-event scenarios. Ablations attribute gains to the DiT backbone, HiST-Sypo priors, and their cross-attention integration.

  • 5.1 SETUP: JavisBench evaluates quality, text consistency, video-audio consistency, and spatio-temporal synchrony across diverse sounding-video scenarios.The benchmark complements existing Landscape and AIST++ evaluations with synchronization-focused metrics.
  • 5.2 MAIN RESULTS AND OBSERVATIONS: JavisDiT achieves strong single-modal quality and semantic alignment, with FVD 204.1, FAD 7.2, TA-IB 0.263, and CLIP similarity 0.302.The authors report these results against UNet-based and naive DiT architectures.
  • 5.2 MAIN RESULTS AND OBSERVATIONS: JavisScore 0.154 surpasses FoleyCrafter for audio-video synchrony on JavisBench.The end-to-end model outperforms cascaded and joint audio-video generation approaches on this synchrony measure.
  • 5.2 MAIN RESULTS AND OBSERVATIONS: FVD 94.2 on Landscape and FAD 9.6 on AIST++ establish state-of-the-art performance on two closed-set benchmarks.The model was trained for 300 epochs under standard settings from prior work.
  • 5.2 MAIN RESULTS AND OBSERVATIONS: Complex multi-object and multi-event videos produce lower JavisScores because they require correct visual-audio correspondence and event-timing modeling.The authors state that current models, including JavisDiT, still struggle with these scenarios.
  • 5.3 IN-DEPTH ANALYSES AND DISCUSSIONS: The full model reaches SAVQ = 6.012, SAVC = 1.201, and SAVS = 0.153 after combining STDiT, HiST-Sypo, and BiCA.HiST-Sypo improves SAVC to 1.191 from 1.155 and SAVS to 0.150 from 0.130, exceeding the gains from simple BiCA.
  • 5.3 IN-DEPTH ANALYSES AND DISCUSSIONS: Increasing the number of spatio-temporal priors consistently improves quality and synchrony, while cross-attention injection outperforms addition and modulation.With 32 priors, alternative injection strategies still improve SAVS to 0.144/0.145 versus 0.133 for the baseline.
  • 5.3 IN-DEPTH ANALYSES AND DISCUSSIONS: JavisDiT maintains visual and auditory quality, semantic consistency, and audio-video synchrony when generating 10-second videos.The reported metrics include FVD, FAD, CLIP, CLAP, AVHScore, and JavisScore.

6 CONCLUSION

The conclusion presents JavisDiT as a unified system for synchronized audio-video generation and JavisBench as a challenging benchmark for real-world evaluation. It also notes practical applications alongside risks from realistic synthetic media and recommends responsible release practices.

  • 6 CONCLUSION: JavisDiT jointly generates high-quality audio and video from open-ended prompts with precise synchronization.Its HiST-Sypo Estimator extracts global and fine-grained priors to guide audio-video alignment.
  • 6 CONCLUSION: JavisBench contains 10,140 text-captioned sounding videos with diverse scenes and real-world complexity.The paper also introduces a temporal-aware semantic alignment mechanism for evaluating complex content.
  • 6 CONCLUSION: JavisBench is constructed from public academic datasets and YouTube videos under stated privacy, licensing, and ethical safeguards.The authors report removing personally identifiable information and retaining no private or user-specific metadata.
  • 6 CONCLUSION: JavisDiT may support creative applications including animation, video conferencing, television broadcasting, and video editing.These domains traditionally rely on offline computation or extensive manual postprocessing for synchronization.
  • 6 CONCLUSION: Realistic temporally coherent synthetic media may be misused for misinformation, impersonation, and manipulation.The paper recommends usage gating, dataset transparency, watermarking, and detection tools as mitigation strategies.

A.5 POTENTIAL LIMITATION AND FUTURE WORK

JavisDiT’s stated limitations concern training-data scale, synchronization-metric accuracy, and computational efficiency. The paper identifies broader and higher-quality data, improved evaluation, and more efficient generation as future directions.

  • Data and scalability: 0.6M training triplets remain limited relative to datasets used by large video-generation foundation models.The paper suggests adding more diverse, higher-quality real-world audio-video samples to improve cross-domain and fine-grained synchronization generalization.
  • Evaluation: JavisScore’s 75% accuracy leaves room for more robust synchronization metrics.Suggested directions include perceptual alignment assessment and human-in-the-loop evaluation.
  • Efficiency: Diffusion-based generation remains computationally intensive despite JavisDiT’s state-of-the-art results.The paper identifies generation speed and efficiency as continuing challenges.
  • Approach: JavisDiT combines shared AV-DiT blocks with spatio-temporal attention and hierarchical synchronized priors for joint generation.Global and fine-grained priors are injected into AV-DiT blocks to guide spatial semantics and temporal synchronization.
  • Prior estimation: Contrastive learning trains the ST-Prior Estimator using synchronous video-audio pairs as positives and asynchronous pairs as negatives.The objective includes token-level hinge, auxiliary discriminative, and VA-embedding discrepancy losses.

C.2.4 NEGATIVE SAMPLE CONSTRUCTION

The benchmark construction targets realistic, diverse synchronization evaluation through multi-dimensional categorization, filtered web videos, and detailed multimodal captions. The resulting JavisBench contains 10,140 curated samples.

  • Taxonomy: Five evaluation dimensions and 19 categories organize JavisBench across scenarios, styles, sounds, subjects, and temporal composition.The hierarchy is used to support evaluation across different aspects of joint audio-video generation.
  • Data sources: YouTube videos uploaded between June and December 2024 supplement existing benchmark test sets while reducing data-leakage risk.The collection combines prior datasets with newly crawled real-world videos.
  • Filtering: Filtering removes unsuitable clips through scene cutting, aesthetic, optical-flow, OCR, and speech checks.Scene segmentation produces 2–60-second clips, while later filters reduce the candidate pool substantially.
  • Annotation: Qwen2-VL, Qwen2-Audio, and Qwen2.5-72B generate, merge, validate, and classify video-audio captions.The pipeline separately captions visual and auditory content before merging them and checking logical conflicts or omissions.
  • Final benchmark: 10,140 samples remain in JavisBench after post-filtering and human checking.The post-filtering removes clips containing only background music or speech voice to improve diversity and balance.

D.4.1 IMPLEMENTATION DETAILS

JavisScore estimates audio-visual synchrony by comparing ImageBind embeddings over overlapping temporal windows. It emphasizes locally desynchronized frames while averaging window scores for stability.

  • Windowed scoring: ImageBind estimates synchrony by comparing video and audio representations within temporally segmented windows.Generated pairs are divided into 2-second windows with 1.5-second overlap before scoring.
  • Local synchronization: The metric focuses on the least synchronized frames so synchronized frames do not dominate the estimate.This makes the score more sensitive to local resynchronization and desynchronization patterns.
  • Global aggregation: JavisScore averages window-level scores across segments to balance local variation and global synchronization.Overlapping windows evaluate frames multiple times, reducing outlier influence and stabilizing the final score.
  • Similarity computation: Cosine similarity between video and audio embeddings underlies the window-level synchrony computation.The formulation compares visual embeddings for sampled video frames with audio embeddings for each segment.
  • Outcome: The resulting approach differentiates synchronized from desynchronized video-audio pairs.This is the stated purpose of the metric’s temporal segmentation and aggregation procedure.

D.4.2 VERIFICATION ON METRIC QUALITY

JavisScore is verified against AV-Align and related metrics using balanced synchronous and asynchronous pairs. It achieves stronger separation, while parameter sensitivity remains limited and perfect accuracy is not reached.

  • Validation design: The 3,000-sample validation set contains 1,000 synchronous pairs and roughly 2,000 asynchronous negatives.Negatives are created through augmentation and additional generated or mismatched sources.
  • Metric comparison: Approximately 0.13 higher AUROC than AV-Align is achieved by JavisScore on 3,000 evaluation samples.The evaluation compares metric scores for positive and negative video-audio pairs.
  • Baseline limitations: AV-Align performs near random guessing in complex scenarios involving subtle visual movements, background noise, or multiple events.Its optical-flow and onset-detection design may miss precise event boundaries.
  • Metric design: JavisScore uses ImageBind’s semantic space to evaluate synchronization at the second level in complex real-world content.The paper presents this as a reason it can distinguish synchronized and unsynchronized cases more robustly than AV-Align.
  • Parameter sensitivity: Approximately 74% accuracy remains the worst result across tested JavisScore parameter settings.The selected configuration uses a 2-second window, 1.5-second overlap, and top 40% minimum strategy, though alternatives do not greatly reduce performance.

D.5 IN-DEPTH ANALYSIS OF EVALUATION RESULTS

JavisBench analysis shows that current joint audio-video generation systems remain weak on rare scenarios, complex compositions, and the relationship between unimodal quality and text consistency. These results expose persistent gaps between research benchmarks and real-world deployment.

  • Unimodal Modeling: Current SOTA models show weak unimodal quality in rare industrial, virtual, animated, ambient, biological, and mechanical scenarios.The analysis attributes some disparities to insufficient training data or overly broad categories.
  • Text Consistency: Unimodal generation quality does not directly correlate with text consistency across visual styles and sound types.Animated scenes can have poor text following despite their unimodal quality, while some sound categories show weak audio-text alignment.
  • Overall Assessment: Current SOTA models still struggle with rare and complex scenarios in both generation quality and audio-video synchronization.The authors identify this as an ongoing challenge for bridging research models and real-world applications.

E.2 EVALUATION ON AUDIO GENERATION QUALITY

The audio-generation evaluation examines training progress and comparisons on AudioCaps and JavisBench-mini, while also analyzing ST-Prior configuration choices. Performance improves with training and is comparable to AudioLDM2 on JavisBench-mini, but AudioCaps remains challenging under variable-length generation.

  • Evaluation Setup: AudioCaps and JavisBench-mini evaluate unimodal audio quality and text following across in-domain and out-of-domain settings.AudioCaps contains 964 filtered samples, while JavisBench-mini contains 1,000 samples used also for ablations.
  • Training Progress: FAD decreases from 5.88 to 5.19 on AudioCaps as training increases from epoch 13 to epoch 55.The same training progression also improves text-consistency measures on JavisBench-mini.
  • Training Progress: TA-IB increases from 0.145 to 0.164 and CLAPScore from 0.368 to 0.381 on JavisBench-mini across training.These changes indicate improved text consistency during the reported training progression.
  • Comparison with AudioLDM2: The model achieves comparable audio-generation performance to AudioLDM2 on JavisBench-mini but performs worse on AudioCaps.The authors attribute the AudioCaps difference partly to AudioLDM2’s explicit 10-second training and the model’s variable-length support.
  • ST-Prior Configuration: A 1:1 spatial-temporal prior token ratio achieves the best performance, while increasing prior number and dimension improves performance with diminishing returns.The selected n32x32+d128 configuration balances performance and training cost.

E.4 TRAINING LOSSES FOR ST-PRIOR’S ESTIMATION

The ST-Prior loss ablations show that different objectives improve synchronization through complementary optimization effects, with the full combination performing best. Additional visualizations illustrate how these priors shift across sequential events and support diverse conditional generation tasks.

  • Loss Design: The ST-Prior Estimator is trained with token-level hinge, auxiliary discriminative, VA-embedding discrepancy, and L2-regularization losses.The ablation evaluates the efficacy of these four objectives.
  • Loss Ablation: Ldisc slightly improves IB-AV from 0.190 to 0.193 when added to Ltoken.The two losses share the goal of separating text-prior anchors from positive and negative video-audio samples but differ in gradient back-propagation.
  • Loss Ablation: Lvad improves JavisScore from 0.136 to 0.140 by increasing divergence between positive and negative video-audio samples.Lreg provides greater benefits through smoother regularization that facilitates convergence toward positive embeddings.
  • Loss Ablation: The complete combination of loss functions achieves the best reported JavisScore of 0.153.The objectives jointly embed video-audio synchrony into the text prior while preserving synchronization semantics.
  • Attention Analysis: During sequential events, visual spatial attention shifts between the heads of the sounding cats, with temporal attention shifting again when the first cat meows again.The example demonstrates dynamic injection of spatiotemporal priors across successive events.
  • Generation and Conditioning: JavisDiT generates synchronized audio-visual pairs across varied environments, styles, audio types, subjects, and temporal compositions.The DiT design also supports conditional tasks such as audio-to-video, video-to-audio, image-to-video, and video-audio extension through dynamic masking.
Loading 2503.23377v2…