Source-linked AI summary

MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance

Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao

arXiv:2608.28212v1cs.MMcs.SD

TL;DR

Models need to reason over both musical scores and performances, but existing benchmarks provide limited coverage of their relationship and of interpretive, long-horizon understanding. MuSP-Bench addresses this gap with 490 human-authored questions across classical piano and orchestral works, evaluated across multiple input conditions. Results show task-dependent strengths: symbolic representations support analytical reasoning, while performance audio and image-based scores remain challenging overall.

  • Problem

    Existing benchmarks provide limited evaluation of integrated score–performance understanding, interpretive listening, and long-horizon musical reasoning.

  • Method

    MuSP-Bench uses 490 questions authored by two professional musicians across 18 classical piano pieces or movements and 6 orchestral excerpts, evaluated with five frontier multimodal models under multiple input conditions.

  • Results

    Model performance is task-dependent: structured symbolic representations are strongest for analytical reasoning, while images and audio help with global stylistic and contextual recognition; score–performance reasoning remains difficult.

  • Takeaways & Limitations

    MuSP-Bench provides a foundation for tracking progress toward meaningful multimodal music understanding across scores, performances, and their relationship.

Abstract

from arXiv · show

Musicians commonly communicate music through scores and performances. Scores encode musical intent, while performances realize it in sound. To investigate whether models can meaningfully engage with both modalities, we introduce MuSP-Bench, a human-authored benchmark of 490 questions targeting understanding across Musical Scores and Performances. The benchmark distinguishes itself by spanning score-based, performance-based, interpretive, and long-horizon reasoning across classical piano and orchestral works. We evaluate frontier multimodal large language models under multiple input conditions. Our results show that these models struggle substantially to understand scores, while facing even greater challenges when reasoning about performance audio. The benchmark is available at https://musp.vaclis.net/.

1. INTRODUCTION

MuSP-Bench addresses a gap in multimodal music evaluation by combining score and performance understanding with interpretive and long-horizon reasoning. Existing benchmarks usually isolate modalities or provide limited coverage of higher-level interpretation and score–performance relationships.

  • Motivation: Scores encode musical intent, while performances realize that intent through time and tone.These representations support musical reasoning in interpretation, imagination, performance, and listening.
  • Existing benchmarks: Existing benchmarks mainly assess score and performance modalities in isolation, often using short or multiple-choice questions.Score benchmarks include ABC excerpts or score images, while several audio benchmarks focus on perception and temporal grounding.
  • Existing benchmarks: Synthetic-score benchmarks provide limited coverage of higher-level musical interpretation.They are useful for symbolic understanding but do not broadly capture interpretive reasoning.
  • Research gap: Existing benchmarks rarely evaluate long-horizon reasoning across score and performance or interpretive listening about climax, phrasing, and technique.MuseBench includes score–audio reasoning but focuses on limited performance attributes and lacks voicing, phrasing, technique, and expression questions.
  • Contribution: MuSP-Bench combines fine-grained music understanding, long-horizon reasoning, and real-world inquiries across score and performance.The benchmark is designed to cover the multimodal reasoning musicians use when engaging with music.

2. BENCHMARK

MuSP-Bench is a human-authored benchmark spanning complete classical piano and orchestral works, multiple modalities, and varied reasoning horizons. Its questions cover analytical, interpretive, and performance-realization concepts across score, audio, MIDI, and their combinations.

  • Benchmark construction: MuSP-Bench contains 490 questions across 18 complete classical piano pieces or movements and 6 orchestral excerpts.Two professional musicians authored the questions.
  • Benchmark construction: 460 questions are open-ended within predefined answer formats, while 30 use multiple options when precise formats were unavailable.Multiple correct answers may be allowed to reflect musical ambiguity.
  • Question dimensions: Questions are annotated by modality as score, performance, jointly score and performance, or either modality.These labels identify the minimum information required to answer each question.
  • Question dimensions: Questions span short, medium, long, and any-horizon evidence requirements, with long questions comprising 36.1% of the benchmark.Short, medium, long, and any categories require different evidence spans from bars or seconds to larger or noncontiguous regions.
  • Content coverage: Coverage ranges from pitch and temporal organization to harmony, form, phrasing, voicing, rubato, articulation, pedaling, style, and work identification.Question actions include identification, localization, quantification, comparison, inference, transcription, explanation, and summarization.

3. EXPERIMENTS

The experiments evaluate five frontier multimodal models under text, score-image, audio, ABC, and MIDI conditions. Inputs are paired with task-relevant modalities, metadata is removed, and performance is measured using normalized answers and a MIDI representation ablation.

  • Models and inputs: Five frontier multimodal models are evaluated using the text, score-image, audio, and combined formats they support.GPT-5.6 Sol and Qwen3.6-Plus use text and score images; Audio Flamingo Next uses audio; Muse Spark 1.2 and Qwen3.5-Omni-Plus support all three.
  • Models and inputs: Score and performance tasks use ABC, PDF score images, audio, and MIDI rendered as text with ABC pitch notation.Score–performance questions combine score images with audio or ABC with audio or MIDI.
  • Controls and evaluation: Metadata is removed from all formats, while a language-only baseline tests answers based only on composer and title.An additional ablation tests how MIDI note representation affects performance.
  • Controls and evaluation: Reported scores use deterministic answer normalization.This standardizes evaluation across the benchmark’s predefined answer formats.

4. RESULTS

Results are strongly task- and representation-dependent: symbolic inputs are strongest for analytical questions, while images and audio help with global stylistic and contextual recognition. Joint score–performance reasoning remains especially difficult, and the reported aggregates weight subsets by question count.

  • Score and performance tasks: 51.1% and 31.7% are the score-only accuracies of GPT and Muse Spark with ABC, respectively.These models perform best on score-only questions with ABC, suggesting stronger processing of structured notation than score images.
  • Score and performance tasks: 59.4% and 30.2% are the performance-only accuracies of GPT and Muse Spark with MIDI-as-text, respectively.Qwen3.5 performs better on audio for performance-only questions.
  • Score and performance tasks: Muse Spark falls from 30.6% with ABC+MIDI to 6.9% with image+audio or ABC+audio on score–performance questions.Qwen3.5 remains below 6% across all three formats.
  • Representation ablation: Replacing ABC pitch names with MIDI note values reduces MIDI-as-text performance by 26.5, 6.3, 0.0, and 7.6 percentage points for GPT, Qwen3.5, Muse Spark, and Qwen3.6, respectively.The comparison shows that note representation materially affects some models’ performance.
  • Aggregate reporting: Table 1 reports accuracy by modality and music representation, with final aggregates weighted by the six subsets’ question counts.Metadata is excluded from the final aggregates.

5. CONCLUSION

MuSP-Bench evaluates interpretive, analytical, and long-horizon reasoning across musical scores and performances. Results show that modality strengths are task-dependent, while performance audio and image-based scores remain challenging overall.

  • MuSP-Bench targets interpretive, analytical, and long-horizon reasoning across musical scores and performances.
  • Structured symbolic representations are strongest for analytical reasoning, whereas images and audio can help with global stylistic and contextual recognition.
  • Performance audio and image-based scores remain challenging overall, so broad-attribute success does not imply robust understanding of musical structure and expression.
Loading 2608.28212v1…