Source-linked AI summary
NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control
Yufan Wen, Zhaocheng Liu, YeGuo Hua, Ziyi Guo, Lihua Zhang, Chun Yuan, Jian Wu
TL;DR
Long-form soundtrack generation must address computational scalability, temporal coherence, and semantic blindness to evolving narrative logic. NarraScore uses frozen VLMs to derive continuous affective trajectories and combines global semantic conditioning with token-level tension control. It achieves state-of-the-art coherence and provides a baseline for autonomous soundtrack generation, while temporal granularity and cascaded error propagation remain limitations.
Problem
Long-form soundtrack generation lacks scalable, temporally coherent modeling of evolving narrative logic, while existing affective priors face domain mismatch and limited data.
Method
NarraScore uses frozen VLMs to extract continuous Valence-Arousal trajectories and injects them through global semantic anchoring and a lightweight token-level affective adapter.
Results
NarraScore achieves state-of-the-art coherence and establishes a strong baseline for autonomous soundtrack generation.
Takeaways & Limitations
The framework provides a direct emotional control pathway from video narratives to musical dynamics using a lightweight adapter and limited labeled data.
Takeaways & Limitations
Affective control lacks temporal granularity for frame-perfect synchronization with rapid visual events, and the cascaded design risks upstream error propagation.
Abstract
from arXiv · showhide
Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to evolving narrative logic. To bridge these gaps, we propose NarraScore, a hierarchical framework predicated on the core insight that emotion serves as a high-density compression of narrative logic. Uniquely, we repurpose frozen Vision-Language Models (VLMs) as continuous affective sensors, distilling high-dimensional visual streams into dense, narrative-aware Valence-Arousal trajectories. Mechanistically, NarraScore employs a Dual-Branch Injection strategy to reconcile global structure with local dynamism: a \textit{Global Semantic Anchor} ensures stylistic stability, while a surgical \textit{Token-Level Affective Adapter} modulates local tension via direct element-wise residual injection. This minimalist design bypasses the bottlenecks of dense attention and architectural cloning, effectively mitigating the overfitting risks associated with data scarcity. Experiments demonstrate that NarraScore achieves state-of-the-art consistency and narrative alignment with negligible computational overhead, establishing a fully autonomous paradigm for long-video soundtrack generation.
1 Introduction
Long-form soundtrack generation must maintain stylistic unity while responding to evolving visual pacing, intensity, and narrative logic. NarraScore addresses this through affective reasoning and separate global and local control.
- Long-form video requires evolving soundtracks that preserve stylistic unity while responding to changing visual pacing and intensity.
- Existing approaches face quadratic memory costs, attention dilution, and style drift when extended from short clips to long-form narratives.
- Surface-level visual representations miss deep narrative cues such as rising tension and resolution, producing semantic blindness.
- NarraScore uses frozen Vision-Language Models to project visual streams into compact, autonomous affective representations and infer evolving emotional arcs from pixels.
- Its Dual-Branch Injection strategy combines a Global Semantic Anchor for stylistic stability with a Token-Level Affective Adapter for local tension modulation.Continuous Valence and Arousal cues are injected as a lightweight additive bias into decoder hidden states.
- NarraScore combines frozen-VLM affective reasoning, data-efficient token-level control, and global-local coordination for automated cinematic soundtrack generation.
2 Related Work
Prior video-to-music systems improve fidelity or acoustic continuity but remain limited by scalability, manual guidance, and weak modeling of evolving narrative semantics. NarraScore targets automated extraction of continuous affective cues aligned with long-form narrative trajectories.
- Early video-to-music methods predicted symbolic MIDI events from visual motion but often required user guidance, limiting expressive diversity and increasing manual dependence.
- Audio-language-model systems project dense video frames into pretrained music backbones, but dense frame-level attention creates scalability bottlenecks for minute-long videos.
- VidMuse and JenBridge improve long-form computational viability through sliding-window inference or video segmentation and stitching.
- These long-form methods can preserve acoustic continuity while overlooking evolving tension, resolution, and other narrative arcs, yielding monotonous ambience.
- Emotion-conditioning methods progressed from discrete frame-level classification and global mapping toward continuous Valence-Arousal representations.
- LLM-based systems improve global emotion labeling but largely avoid continuous emotion curves because large-scale continuous affective data are scarce.
- Fine-grained methods relying on extrinsic guidance remain difficult to scale autonomously, leaving parameter-efficient extraction of continuous affective cues as a critical gap.
- Off-the-shelf affective recognizers are mismatched to soundtrack generation because face-centric benchmarks and noisy predictions may not capture induced atmosphere or narrative tension.
3 Methodology
NarraScore decomposes long-form video-to-music generation into global atmospheric modeling and local tension tracking, using frozen VLM reasoning to derive affective controls. These controls condition a pretrained acoustic decoder through semantic anchoring and selective token-level injection.
- Problem Definition: The task requires long-form soundtracks to maintain unified musical style while synchronizing with frame-level narrative progression.The acoustic sequence is dense, whereas the visual sequence remains sparse to respect minute-level processing limits.
- Overview: NarraScore separates visual-to-music generation into a global semantic anchor and a frame-level affective trajectory.The two priors represent macro-scale atmospheric modeling and micro-scale tension tracking, respectively.
- Narrative-Aware Affective Reasoning: A frozen Vision-Language Model is probed with aligned visual sequences and semantic instructions to extract narrative-aware affective representations.Temporal snapshots are interleaved with textual semantic clocks, while the instruction primer suppresses low-level object enumeration and promotes narrative reasoning.
- Narrative-Aware Affective Reasoning: A lightweight probing head aggregates frame visual tokens and projects them onto the Valence-Arousal plane to produce the local affective trajectory.Only the lightweight probe is trained, while the massive backbone remains frozen, reducing computational overhead.
- Hierarchical Acoustic Synthesis: The global semantic anchor conditions the acoustic decoder toward the target genre and atmosphere through pretrained cross-attention.This conditioning guides musical structure and instrumentation while preserving the pretrained backbone's generative distribution.
- Hierarchical Acoustic Synthesis: A temporal super-resolution adapter maps sparse Valence-Arousal cues into dense acoustic controls, which are injected additively into shallow Transformer blocks.The adapter addresses the resolution mismatch between sparse visual emotion cues and dense acoustic tokens; the control signal is applied through residual modulation.
4 Experiments
Experiments evaluate NarraScore using affective video and music datasets, standardized preprocessing, trainable feature-alignment modules, objective metrics, and subjective comparisons against five baselines. Results report superior performance across metrics, stronger long-form preference, and benefits from explicit affective reasoning.
- Datasets: The study trains video-to-emotion and emotion-to-music modules on publicly available continuous-affective benchmark datasets.The video dataset provides frame-level induced Valence-Arousal annotations, while the music dataset contains dense per-second affective labels.
- Datasets: Music emotion captions combine genre and tags with dynamic Valence-Arousal labels to decouple emotional control from instrumentation-specific style.Continuous affective shifts are associated with changes in harmony, tempo, and instrumentation.
- Implementation Details: The framework uses VideoLlama-3 and MusicGen-Small with a Projector for visual-to-acoustic alignment and a Temporal Adapter for long-term affective dependencies.The Projector is a two-layer MLP with GELU and 0.1 dropout; the adapter uses dilated convolution, LeakyReLU, and a linear layer.
- Implementation Details: Training first aligns static visual and acoustic features for 150 epochs, then fine-tunes the Adapter for 50 epochs to capture temporal dynamics.
- Evaluation: NarraScore is evaluated against M2UGEN, video2music, VidMuse, GVMGEN, and Caption2Music using objective, diversity, and cross-modal consistency metrics.Metrics include FAD, FD, KLD, Density, Coverage, and ImageBind score.
- Results: NarraScore leads across subjective metrics, achieves the highest overall preference in short/mid- and long-form settings, and widens its margin for long-form videos.The reported advantages include emotional consistency, stable global musical identity, and adaptive response to local visual cues.
- Results: Ablations show that NAR consistently improves music quality across backbone configurations, while a 75% injection ratio balances narrative guidance and acoustic modeling.Both lower and higher injection proportions adversely affect performance.
- Qualitative Analysis: Qualitative spectrogram analysis links Caption2Music to smooth bands, VidMuse to high-frequency noise above 8000Hz, and NarraScore to harmonic stability with rhythmic precision.NarraScore’s spectral structure is described as synchronizing rhythmic events with the narrative arc while retaining acoustic clarity.
5 Conclusion
NarraScore establishes a direct emotional control pathway from video narratives to musical dynamics. The framework combines small-model fine-tuning and lightweight adaptation to achieve robust regression and state-of-the-art coherence for autonomous soundtrack generation.
- Limited labeled data enables robust continuous temporal regression in NarraScore.
- A lightweight adapter sufficiently steers the complex acoustic backbone while reconciling global style with local tension.
- NarraScore achieves state-of-the-art coherence as a baseline for autonomous soundtrack generation.
6 Limitations
NarraScore’s affective control has limited temporal granularity, restricting synchronization with rapid visual events and leaving cascaded error propagation as an additional risk.
- Limited temporal granularity precludes frame-perfect synchronization with rapid visual events.
- The cascaded design risks error propagation from upstream affective reasoning.