Source-linked AI summary
Audio-Visual Intelligence in Large Foundation Models
You Qin, Kai Liu, Shengqiong Wu, Kai Wang, Shijian Deng, Yapeng Tian, Junbin Xiao, Yazhou Xing, Yinghao Ma, Bobo Li, Roger Zimmermann, Lei Cui, Furu Wei, Jiebo Luo, Hao Fei
TL;DR
Audio-Visual Intelligence lacks a unified account across its diverse tasks, methods, and evaluation practices, limiting systematic comparison. This survey organizes AVI through a taxonomy and synthesis of representations, models, resources, and future directions, concluding that cross-modal alignment remains foundational while data and faithful context memory remain key boundaries.
Problem
AVI research is fragmented across tasks, taxonomies, terminology, and evaluation practices, despite the need to model synchronized audio and vision for multimodal understanding, generation, and interaction.
Method
The survey unifies AVI around understanding, generation, and interaction while synthesizing tokenization, fusion, autoregressive and diffusion modeling, alignment, benchmarks, and context-memory designs.
Results
The review finds that cross-modal alignment remains the foundational challenge across AVI tasks and organizes the field into representation-, generation-, and LLM-centric methods.
Takeaways & Limitations
Future AVI research is framed around causal, contextual, controllable, verifiable, and interactive systems supported by structured memory and evaluation of delayed recall, audio-critical questions, modality conflict, and belief revision.
Takeaways & Limitations
Omni-modal conversation remains constrained by insufficient naturalistic audio-visual dialogue data, because synthetic instruction data rarely captures authentic timing, grounded events, or human spontaneity.
Abstract
from arXiv · showhide
Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate, and interact in the multimodal real world. In the era of large foundation models, joint modeling of audio and vision has become increasingly crucial, i.e., not only for understanding but also for controllable generation and reasoning across dynamic, temporally grounded signals. Recent advances, such as Meta MovieGen and Google Veo-3, highlight the growing industrial and academic focus on unified audio-vision architectures that learn from massive multimodal data. However, despite rapid progress, the literature remains fragmented, spanning diverse tasks, inconsistent taxonomies, and heterogeneous evaluation practices that impede systematic comparison and knowledge integration. This survey provides the first comprehensive review of AVI through the lens of large foundation models. We establish a unified taxonomy covering the broad landscape of AVI tasks, ranging from understanding (e.g., speech recognition, sound localization) to generation (e.g., audio-driven video synthesis, video-to-audio) and interaction (e.g., dialogue, embodied, or agentic interfaces). We synthesize methodological foundations, including modality tokenization, cross-modal fusion, autoregressive and diffusion-based generation, large-scale pretraining, instruction alignment, and preference optimization. Furthermore, we curate representative datasets, benchmarks, and evaluation metrics, offering a structured comparison across task families and identifying open challenges in synchronization, spatial reasoning, controllability, and safety. By consolidating this rapidly expanding field into a coherent framework, this survey aims to serve as a foundational reference for future research on large-scale AVI.
1 Introduction
Audio-Visual Intelligence unifies audio and vision for understanding, controllable generation, and interactive reasoning, but its rapidly expanding literature remains fragmented. This survey organizes the field through a unified taxonomy, methodological synthesis, and structured coverage of resources and challenges.
- Task Landscape: AVI spans perception, generation, and interaction tasks that require synchronization and grounding across audio and vision.Representative tasks include speech recognition, sound localization, audio-driven talking heads, video-conditioned sound, dialogue, and embodied interaction.
- Motivation: The literature is fragmented by overlapping definitions, inconsistent terminology, divergent taxonomies, and heterogeneous evaluation practices.Variation is especially pronounced for open-ended generation, alignment quality, temporal coherence, and human-centric judgments.
- Contributions: The survey provides a first systematic review of AVI within the large-foundation-model paradigm, unifying perception, generation, and interaction.It frames the field as a coherent research landscape rather than a collection of isolated subcommunities.
- Contributions: A principled taxonomy organizes audio-visual tasks across speech, music, sound events, video, and open-world understanding, generation, and interaction.The taxonomy clarifies task scope, assumptions, and relationships among subproblems.
- Methodological Foundations: The methodological synthesis covers modality tokenization, cross-modal fusion, autoregressive and diffusion generation, large-scale pretraining, and instruction alignment.These foundations support sequence modeling, synthesis, conditional control, and interactive use.
- Resources and Frontiers: The survey curates datasets, benchmarks, and metrics, identifies assessment gaps, and highlights synchronization, spatial reasoning, controllability, safety, watermarking, and governance challenges.It also reports plans to publicly release summarized resources, references, and organizational structures.
2 Preliminary
The preliminary section defines audio and visual data in digital form and surveys the representation families used by audio-visual foundation systems. It connects these representations to multimodal encoders, tokenizers, generative models, and unified interaction systems.
- Audio Modality: Audio includes speech, music, and general sound events, while digital audio is commonly represented as mono or stereo waveform time series.Multi-channel spatial audio formats are also possible.
- Audio Representation: Audio signals are primarily modeled as raw waveforms or time-frequency representations such as log-Mel spectrograms.Waveforms can be mono, stereo, or multi-channel; spectrograms derive from Short-Time Fourier Transform processing.
- Representation Families: Audio-visual systems use global embeddings, continuous dense representations, and discrete token sequences to encode modality-specific information.Global embeddings capture segment-level semantics, dense features preserve temporal or frequency structure, and discrete tokens support language-model-style modeling.
- Encoders: Encoders map audio and visual inputs into lower-dimensional embeddings or feature maps that feed multimodal and generative systems.Audio inputs may be waveforms or time-frequency transforms, while visual encoders produce downsampled feature maps or sequences.
3 Task Taxonomy
Audio-Visual Intelligence is organized across understanding, generation, and interaction, with perception itself progressing from low-level sensing to semantic understanding and logical reasoning. The taxonomy also distinguishes conditional, cross-modal, and joint generation, alongside conversational and embodied interaction.
- AVI tasks span understanding, generation, and interaction as a roadmap for developing general audio-visual intelligence.
- Understanding: Perception progresses from pixel-level sensing and alignment to content understanding and logical reasoning over causes and outcomes.
- Generation: Generation is categorized into conditional, cross-modal, and joint paradigms with increasing integration of audio and vision.
- Interaction: Interactive audio-visual systems combine ongoing multimodal interpretation with reasoning and timely outputs or actions.
- Interaction: Interactive conversation supports audio-centric, visual-centric, and omni-modal interfaces, while embodiment covers navigation, embodied question answering, and manipulation.
4 Foundation Techniques
Foundation techniques for AVI comprise representation-centric, generation-centric, and LLM-centric methods. They use self-supervised alignment, continuous or discrete tokenization, multimodal generation mechanisms, and integrated architectures while facing efficiency, synchronization, and tokenization challenges.
- AVI foundation techniques are organized into representation-centric, generation-centric, and LLM-centric methods.
- Representation: Representation methods transform audio and visual signals into continuous embeddings or discrete tokens for neural processing.
- Representation: Self-supervised, contrastive, correlation-based, and task-oriented methods learn shared audio-visual representations from synchronization and cross-modal interactions.
- Representation: Multimodal VAEs learn structured shared and modality-specific latent spaces that support cross-modal reconstruction and generation.
- Generation: Generation-centric methods cover conditional synthesis and editing across unimodal and cross-modal tasks, with semantic coherence and temporal synchronization as key requirements.
- Generation: GANs, diffusion and continuous-time models, and autoregressive transformers provide alternative generation mechanisms for audio-visual synthesis.
- LLM-centric methods: Unified multimodal systems can reduce latency and brittle interfaces, but heterogeneous tokenization, long-context modeling, streaming, and open-ended evaluation remain challenges.
5 Audio-Visual Perception
Audio-visual perception follows a progression from raw signal structure to contextual meaning and then to relations among events over time. The section distinguishes unimodal previews from audio-visual problems while organizing perception along this abstraction ladder.
- Perception moves from pixel- and sample-level structure to semantic content understanding and reasoning beyond one-step recognition.
5.1 Audio-Visual Pixel Perception
Audio-visual pixel perception covers low-level structure and alignment in sound and vision, progressing toward locating jointly audible and visible events. Recent work increasingly builds on foundation models, audio-guided prompting, reasoning, and generation rather than dense pixel-level fusion alone.
- Audio Pixel Perception: Audio perception extracts low- and mid-level acoustic structure, including event presence, temporal boundaries, onsets, beats, voice activity, and separated sources.Representative methods include self-supervised encoders for reusable representations, temporal labeling models, and source-separation systems.
- Visual Pixel Perception: Visual pixel perception predicts spatially resolved outputs for pixels or regions to ground objects, actions, and events across space and time.The section highlights detection, segmentation, tracking, and temporal action detection as visual primitives for audio-visual integration.
- Audio-Visual Event Localization: Audio-visual event localization identifies when and where events are simultaneously present in audio and video, using cross-modal attention, co-attention, gated attention, and adaptive fusion.Pre-trained audio and video encoders and lightweight cross-modal adapters are also used to improve robustness and parameter efficiency.
- Emerging Directions: Overall, audio-visual segmentation is shifting from pixel-level fusion toward audio-guided prompting, reasoning, and generation on foundation models, reducing reliance on dense annotations.Related trends include reasoning-centric segmentation and more stable synchronization formulations based on offset classification.
- Emerging Directions: Audio-visual synchronization has progressed from lip-centric alignment to general cross-modal temporal correspondence learning, including scalable and perceptually grounded offset classification.Future directions emphasize robustness, finer temporal dynamics, foundation models, and integration with broader multimodal systems.
5.2 Audio-Visual Content Understanding
Audio-visual content understanding maps audio and visual signals to semantic interpretations of objects, events, questions, and retrieval targets. The field is moving toward language-centric foundation models with finer temporal grounding, evidence alignment, and modality-aware reasoning.
- Audio Understanding: Audio content understanding interprets acoustic signals semantically through speaker recognition, emotion analysis, audio captioning, and audio question answering.These tasks range from identifying speaker traits and affective states to generating descriptions and answering questions grounded in audio.
- Visual Understanding: Visual content understanding extracts objects, attributes, actions, and spatial relations from images and videos into structured or textual interpretations.Modern systems couple visual encoders with language models for recognition, grounding, captioning, and video question answering.
- Audio-Visual Question Answering: Audio-visual question answering jointly relates sounds, visual events, language, temporal dynamics, and sound sources to answer natural-language questions.Methods improve question-conditioned clue extraction, cross-modal alignment, temporal modeling, and distractor suppression.
- Audio-Visual Question Answering: AVQA is shifting from task-specific fusion models to MLLM-based, language-centric reasoning centered on fine-grained alignment, evidence grounding, bias mitigation, and efficient inference.This shift emphasizes whether models truly reason over audio-visual evidence rather than relying only on raw accuracy.
- Cross-Modal Retrieval: Audio-visual retrieval is moving from pairwise alignment toward scalable multimodal representation learning and finer, temporally grounded retrieval.Contrastive and masked modeling with large-scale pretraining are reported as improving alignment and generalization.
5.3 Audio-Visual Logical Reasoning
Audio-visual logical reasoning extends perception and question answering to multi-step, temporal, causal, and grounded inference over sound and video. Recent systems increasingly use multimodal LLMs and reasoning-oriented training, but faithful cross-modal reasoning and reliable evaluation remain limited.
- Audio Reasoning: Audio reasoning adapts LLM-style prompting, instruction tuning, and reasoning-oriented training to acoustic evidence, although complex temporal and causal relationships remain difficult.Audio-specific prompts can separate perception from inference, while instruction data and multitask supervision target broader reasoning generalization.
- Visual Reasoning: Visual reasoning concerns multi-step, compositional, and abstract inference involving relations, counting, causality, and commonsense grounded in visual inputs.The literature has progressed from symbolic and neuro-symbolic task-specific models to vision-language foundation models and reinforcement-based fine-tuning.
- Audio-Visual Reasoning: Audio-visual reasoning infers why, when, and how events unfold by jointly interpreting video and audio with temporal alignment, causal inference, and grounded rationales.It goes beyond recognition and simple question answering by requiring multi-step inference over cross-modal evidence.
- Audio-Visual Reasoning: Industrial and open omni-modal systems increasingly position audio-visual understanding as reasoning-centric, extending multimodal perception toward explicit thinking over audio and video.Examples include OpenAI-o3, Gemini 2.5/3.0 Pro, and Qwen3-Omni-Thinking.
- Open Challenges: Audio-visual reasoning remains at an early stage, with limitations in faithful cross-modal reasoning, robust temporal understanding, and reliable evaluation.Reported concerns also include hallucinations, spurious correlations, unfaithful reasoning, and insufficient benchmarks for correctness and faithfulness.
6 Audio-Visual Generation
Audio-visual generation synthesizes temporally aligned sound and imagery from text, images, video, or audio. The field progresses from strong unimodal generators to cross-modal translation and increasingly coupled or unified architectures.
- Task Organization: Audio-visual generation covers conditional unimodal generation, cross-modal translation, and joint audio-visual generation.These regimes organize synthesis from text, images, video, or audio while emphasizing temporal alignment between sound and imagery.
- Architectural Evolution: The field has moved from composing strong single-modality generators toward explicitly coupled and more recently unified architectures.Diffusion, flow matching, and multimodal transformers are identified as key components of this progression.
6.1 Conditional Audio/Visual Generation
Conditional generation synthesizes or transforms audio and visual content under external control, while serving as a foundation for cross-modal and joint audio-visual generation. The section covers text, reference-media, structured-control, and instruction-guided regimes across audio and visual tasks.
- Conditional Audio Generation: Conditional audio generation synthesizes or transforms sound under text, reference audio, mixture recordings, or editing instructions.
- Conditional Audio Generation: Text-conditioned systems generate sound effects, ambient scenes, music, or speech using diffusion, codec-token language models, and broader conditional frameworks.Neural audio tokenizers and latent representations support fidelity, controllability, and efficiency.
- Conditional Audio Generation: Conditional audio transformation covers enhancement, source separation, remixing, restoration, voice conversion, inpainting, and style transfer.Promptable and frequency-structured methods support open-set separation and instrument-level control.
- Conditional Visual Generation: Conditional visual generation synthesizes or transforms images and videos under text, reference images, structured controls, or editing instructions.
- Conditional Visual Generation: Text-to-image and text-to-video systems have progressed from early generative models toward stronger multimodal foundation and product systems.
- Conditional Visual Generation: Controllable video generation separates appearance, structure, motion, and temporal continuation, while image and video editing seeks semantic changes that preserve scene identity and temporal coherence.Pretrained diffusion models absorb structured controls, and instruction-based methods extend from image editing to video workflows.
6.2 Audio-Visual Cross-Modal Generation
Cross-modal generation drives one sensory stream from another, requiring semantic correspondence, temporal alignment, and controllability. The section covers video-to-audio, audio-to-video, and audio-synchronized image animation, with increasing emphasis on explicit reasoning and long-horizon generation.
- Video-to-Audio Generation: Video-to-audio generation adds Foley, ambience, speech, or music to silent video while satisfying semantic relevance, event synchronization, and acoustic plausibility.
- Video-to-Audio Generation: Video-to-audio methods include alignment-first diffusion or flow models, jointly trained controllable systems, and reasoning-guided approaches for open-world and long-form synthesis.These methods use contrastive pretraining, onset conditioning, multimodal joint training, text-guided control, and higher-level scene reasoning.
- Video-to-Audio Generation: Recent video-to-audio systems increasingly infer events, sources, and timing before rendering sound, making multimodal reasoning a central design component.The field is shifting from direct video-feature mapping toward structured audiovisual planning followed by audio generation.
- Audio-to-Video Generation: Audio-to-video methods use adapted video backbones, motion-mediated pipelines, or identity-conditioned generation to improve alignment, control, or appearance preservation.Motion-mediated systems strengthen beat alignment but add pipeline complexity and weaken end-to-end flexibility.
- Audio-to-Video Generation: Audio-to-video generation is underdetermined because one soundtrack can support many plausible videos, creating a trade-off between diversity and user steerability.Reference images, motion priors, and explicit controls reduce ambiguity but do not fully resolve this trade-off.
- Audio-Synchronized Image Animation: Audio-synchronized image animation generates video from a reference image and driving audio while preserving identity and temporal coherence.Recent systems extend beyond talking heads toward semi-body, full-body, multi-character, and general image-conditioned animation.
6.3 Joint Audio-Visual Generation
Joint audio-visual generation produces or edits coupled sound and video within shared objectives or tightly integrated workflows. The literature advances unified architectures, synchronization strategies, multi-task training, agentic planning, and more precise localized editing, while scale and benchmarking remain limiting factors.
- Joint Audio-Visual Generation: Joint generation models the coupled evolution of sound and vision rather than attaching a unimodal output after generation.
- Text-to-Audio-Video Generation: Text-to-audio-video generation must jointly satisfy within-modality quality, prompt faithfulness, and cross-modal synchronization.
- Text-to-Audio-Video Generation: Joint audio-video architectures have evolved from partially separate branches with synchronization modules toward shared or tightly coupled attention blocks.Aligned temporal priors, frame-level cross-attention, aligned position identifiers, and data scaling support synchronization.
- Text-to-Audio-Video Generation: Multi-task systems integrate A2V, V2A, and image-conditioned tasks to improve modality learning and cross-task synergy, with some work targeting real-time generation.
- Text-to-Audio-Video Generation: Agentic workflows decompose audio-visual generation into cascaded stages, typically text-to-video followed by video-to-audio, with explicit planning for long-form generation.
- Text-to-Audio-Video Generation: Commercial systems remain ahead through larger audio-video corpora, stronger base generators, and heavier post-training, while open models show promising performance and scalability.The remaining gap is attributed to combined data scale, base-model maturity, and post-training depth rather than one architectural trick.
- Joint Audio-Video Editing: Joint audio-video editing is shifting from coarse modification toward local, synchronous intervention with minimal side effects.Newer benchmarks add masks or paired source-target supervision to evaluate edit faithfulness, preservation, and synchronization.
7 Audio-Visual Interaction
Audio-visual interaction requires real-time responses grounded in multimodal perception, memory, reasoning, and generation. Conversational systems are moving toward richer acoustic and omni-modal interfaces, while unified video modeling and naturalistic dialogue data remain immature.
- Audio-Visual Interaction: Audio-visual interaction includes conversational responses and embodied physical actions, both requiring tightly coupled multimodal processing under latency constraints.This makes interaction a stricter test of audio-visual intelligence than static understanding alone.
- Audio-Driven Conversation: Audio-driven conversation accepts speech, environmental sound, or music and responds in text or speech while preserving acoustic, linguistic, and paralinguistic information.
- Audio-Driven Conversation: Audio conversation models use unified or domain-specific encoders to preserve emotion, timbre, speaker identity, and environmental sound.
- Audio-Driven Conversation: Speech interaction is moving beyond hard ASR→LLM→TTS cascades, but cascades retain advantages in modularity and controllable intermediate text.End-to-end systems improve latency and prosody preservation, while token-based hybrids offer a practical compromise.
- Unified Visual Understanding and Generation: Unified visual understanding and generation seeks one model for interpreting and synthesizing images or videos within a shared conversational loop.
- Unified Visual Understanding and Generation: Unified video modeling remains early, with most high-performing systems using hybrid MLLM-generator designs because visual quality gaps persist.Continuous or latent-token autoregression is emerging as a middle ground between discrete unification and hybrid systems.
- Omni-Modal Conversation: Omni-modal conversation remains constrained by limited naturalistic data capturing conversational timing, grounded audio events, and human spontaneity.Synthetic instruction data scales but does not fully capture these properties.
- Embodied Interaction: Embodied audio-visual agents need spatially consistent representations that support movement decisions, out-of-view reasoning, and predicting how environments sound after actions.Map-like spatial memory, queryable acoustic fields, and world-model interfaces are emphasized over standalone classification accuracy.
8 Applications
AVI foundation models support applications across creative production, communication, accessibility, immersive environments, and embodied systems. These applications combine audio-visual understanding, generation, and interaction in increasingly integrated workflows.
- Creative Industries: AVI foundation models are reshaping film, video, music, and sound-design workflows through generation, editing, Foley synthesis, and source separation.Video-to-audio models automate temporally aligned sound effects, while text-to-audio and separation systems support rapid prototyping and remixing.
- Avatars and Communication: Audio-driven avatars have progressed from pixel-level lip synchronization to photorealistic 3D talking heads and identity-consistent personas.Applications include virtual influencers, customer-service avatars, and personalized AI companions.
- Human-Centered Applications: Omni assistants, educational systems, and accessibility tools use audio-visual signals for multimodal conversation, adaptive instruction, synchronized narration, and sensory support.Audio-native models preserve paralinguistic cues, while captioning and audio description improve access for users with sensory impairments.
- Immersive Technologies: Immersive applications jointly model sound, images, and spatial context to support movement-responsive rendering, scene queries, and 3D interaction.Neural acoustic fields and joint audio-visual representations enable spatial audio, while acoustic reflections can supplement visual sensing when line of sight is limited.
- Embodied Systems: Embodied agents combine vision and sound for navigation, while contact audio provides information about materials, slips, and grasp stability during manipulation.Audio-visual navigation uses reverberation, intensity gradients, and visual context to locate sound-emitting targets.
9 Open Challenges, Limitations, and Future Directions
AVI’s central challenge is modeling audio and vision as partially observed views of a dynamic world rather than merely matching correlated signals. The survey proposes a roadmap toward causal, contextual, controllable, verifiable, interactive, and responsible systems.
- Core Challenge: AVI must explain why audio-visual evidence co-occurs and how events, propagation, materials, occlusions, and conversational context shape those observations.This distinguishes AVI from generic multimodal learning and motivates causal grounding beyond correspondence.
- Roadmap: The proposed roadmap advances from correspondence, perception, and generation toward interactive systems, then causal-contextual and verifiable agentic AVI.Table 24 links six research axes to concrete limitations and required structural capabilities.
- Causal Grounding: Temporal synchronization checks local agreement, whereas causal event-source grounding identifies which source produced each sound through which causal path.The agenda includes event-source graphs and counterfactual supervision for provenance, timing, propagation, and uncertainty.
- World Models: Audio-visual world models should forecast compact action-conditioned latent states covering sources, materials, acoustics, affordances, intent, and uncertainty.The proposed abstraction supports counterfactual queries and combines learned dynamics with approximate acoustic physics and task-scoped scorers.
- Evaluation: Evaluation should prioritize navigation, manipulation, spatial reasoning, and counterfactual prediction rather than waveform or pixel reconstruction alone.Rendering-centric fields, navigation benchmarks, and pixel-heavy generators do not yet provide a shared predictive abstraction.
- Context Memory: Long multimodal contexts require layered audio-visual memory that preserves raw evidence, event structure, semantic summaries, and user or task state.OmniVideoBench duration splits show that high short-video accuracy does not automatically yield robust ultra-long-video accuracy.
- Controllability: Causal editing should represent objects, stems, identities, motions, and dependencies so interventions freeze untouched content while propagating necessary changes.This addresses prompt entanglement and edits that leak unrelated ambience or fail to update counterpart modalities.
- Interactive and Responsible AVI: The frontier combines causal grounding, world models, faithful memory, causal intervention, multimodal verifiers, and responsible interaction for safe deployment alongside humans.These capabilities extend AVI beyond larger models and datasets toward systems that explain evidence, forecast consequences, and close the loop with verification.
10 Conclusion
The survey organizes foundation-model AVI around representation-centric embedding, generation-centric cross-modal synthesis, and LLM-centric reasoning. It traces progress from event recognition to complex spatial reasoning.
- Conclusion: The survey structures AVI around representation-centric methods, generation-centric cross-modal synthesis, and LLM-centric reasoning engines.It also traces audio-visual understanding from event recognition toward complex spatial reasoning.