Source-linked AI summary

Creating A Multi-track Classical Musical Performance Dataset for Multimodal Music Analysis: Challenges, Insights, and Applications

Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, Gaurav Sharma

arXiv:1612.08727v3cs.MMcs.SD

TL;DR

Multimodal music analysis lacks datasets combining synchronized isolated tracks, visual recordings, scores, and detailed annotations. This paper constructs URMP from 44 classical ensemble pieces using a conducting-video approach, with video-based vibrato analysis achieving generally over 90% F-measure and low parameter-estimation errors. The dataset supports established and emerging MIR tasks, but its single-camera visual setup limits some applications.

  • Problem

    Multimodal music-performance analysis is constrained by scarce datasets containing recordings, visual data, and ground-truth annotations.

  • Method

    The paper constructs URMP from classical ensemble pieces with separately recorded tracks and uses a pre-recorded conducting video to coordinate individual performers.

  • Results

    Video-based vibrato detection generally exceeds 90% F-measure, with average estimation errors of 0.38 Hz for rate and 3.47 musical cents for extent.

  • Takeaways & Limitations

    URMP supports established MIR benchmarks and emerging multimodal tasks including polyphonic vibrato analysis, visually informed source separation, and cross-modality generation.

  • Takeaways & Limitations

    URMP uses a single camera with sometimes non-optimal views, limiting visual coverage and making some applications such as depth estimation or 3D reconstruction unsupported.

Abstract

from arXiv · show

We introduce a dataset for facilitating audio-visual analysis of music performances. The dataset comprises 44 simple multi-instrument classical music pieces assembled from coordinated but separately recorded performances of individual tracks. For each piece, we provide the musical score in MIDI format, the audio recordings of the individual tracks, the audio and video recording of the assembled mixture, and ground-truth annotation files including frame-level and note-level transcriptions. We describe our methodology for the creation of the dataset, particularly highlighting our approaches for addressing the challenges involved in maintaining synchronization and expressiveness. We demonstrate the high quality of synchronization achieved with our proposed approach by comparing the dataset with existing widely-used music audio datasets. We anticipate that the dataset will be useful for the development and evaluation of existing music information retrieval (MIR) tasks, as well as for novel multi-modal tasks. We benchmark two existing MIR tasks (multi-pitch analysis and score-informed source separation) on the dataset and compare with other existing music audio datasets. Additionally, we consider two novel multi-modal MIR tasks (visually informed multi-pitch analysis and polyphonic vibrato analysis) enabled by the dataset and provide evaluation measures and baseline systems for future comparisons (from our recent work). Finally, we propose several emerging research directions that the dataset enables.

I. INTRODUCTION

Audio-visual music performance analysis has expanded, but progress remains constrained by scarce datasets combining recordings, visual data, and ground-truth annotations. The URMP dataset addresses this gap with coordinated multi-track classical performances designed for existing and novel MIR tasks.

  • Visual information contributes to music perception, learning, performance analysis, transcription, onset detection, and vibrato analysis.Examples include tracking violin strings and fingers and analyzing piano, guitar, and drum performances.
  • Progress in joint audio-visual music analysis has been slow partly because suitable datasets and ground-truth annotations are scarce.Creating annotations is time consuming, often requires musical expertise, and source-separation research also needs isolated recordings.
  • URMP contains 44 classical chamber pieces with scores, isolated instrument tracks, assembled audio and video, and frame-level and note-level transcriptions.The pieces range from duets to quintets and were assembled from separately recorded instrumental sources.
  • URMP targets both established audio-only MIR tasks and novel multimodal tasks enabled by its audio-visual multi-track design.The paper benchmarks existing tasks and identifies additional multimodal research directions.

II. REVIEW OF MUSIC PERFORMANCE DATASETS

Existing music performance datasets support different MIR needs, but most are audio-only and vary substantially in track structure, annotations, scale, and modality. Multi-track datasets offer versatility, while multimodal and classical ensemble coverage remains limited.

  • Music datasets are difficult to create because recording performances and producing expert ground-truth labels are time consuming, while commercial recordings face copyright constraints.Laboratory collection also depends on available musicians and recording facilities.
  • Most reviewed datasets contain only audio, and only six are multi-modal.This limits the availability of paired visual and musical-performance data for multimodal analysis.
  • Single-track polyphonic datasets commonly provide MIDI transcriptions, but accurate transcription for non-MIDI instruments often requires labor-intensive manual annotation.Some datasets instead align reference MIDI scores to audio performances semi-automatically.
  • Multi-track datasets support source separation, reduce annotation complexity through monophonic tracks, and allow varied mixtures from the same piece.Their principal recording difficulty is synchronizing isolated instrumental sources.
  • Existing multi-track collections span large pop and rock datasets, small classical ensembles, and datasets with MIDI, pitch, note, or alignment annotations.Examples include MedleyDB, SSMD, iKala, WWQ, TRIOS, Bach10, and PHENICX-Anechoic.
  • Existing multimodal datasets include guitar, clarinet, drums, instrument-playing, and ensemble-performance recordings, with differing visual and audio annotations.The reviewed examples include both single-track and multimodal performance datasets.

B. Synchronization Challenges in Creating Multi-track Datasets

Isolated recording avoids cross-track leakage but removes the live interaction that normally supports synchronization. The authors compare increasingly expressive synchronization strategies and identify a conducting-video procedure as a practical compromise, while noting limits of onset-deviation metrics.

  • Separate recording is needed to eliminate leakage, but players cannot use live interactions to adjust timing across instrumental parts.This creates a central synchronization challenge for realistic multi-track datasets.
  • Classical ensembles are especially difficult to synchronize because their performances vary in tempo and dynamics and are usually recorded together.Approaches developed for steady-tempo popular music do not directly transfer to classical performance.
  • A pre-recorded piano performance improved expressiveness over a metronome but produced unsatisfactory synchronization and outliers after long breaks.Players lacked sufficient cues for starting or re-entering after rests.
  • Joint rehearsal with a conductor improved synchronization to a median onset microtiming value of 16ms but was time-consuming and difficult to scale.The approach used rehearsal recordings as synchronization references.
  • Following a professional performance video was difficult for players, preventing completion and quantitative synchronization analysis in that attempt.The visual cues were reported to be unclear even after repeated viewing.
  • The final approach used a conducting video with pianist audio and cues for instrumental entries, balancing synchronization, expressiveness, and recording burden.The conductor led the pianist and signaled parts before entries after long rests.
  • Onset-time deviation is only a limited synchronization indicator because onset times are ambiguous for some soft articulations.The authors therefore also considered players’ subjective evaluations.

IV. DATASET CREATION PROCEDURE

URMP’s creation procedure covered selection, recruitment, recording, post-production, and annotation. Piece selection prioritized broad instrumentation and polyphony, manageable complexity, and sufficient room for expressive interpretation.

  • The dataset-creation workflow spans piece selection, musician recruitment, recording, post-production, and ground-truth annotation.The complete process is summarized using a duet example.
  • A. Piece Selection: Piece selection sought coverage of polyphony, composers, and instrumentations across duets, trios, quartets, and quintets.Selected instrumentation included string, woodwind, brass, and mixed groups, excluding percussion.
  • A. Piece Selection: Pieces were kept relatively simple and short so players could perform them with limited practice and recording demands.The target duration was ideally one to two minutes, and many players could sight-read or practice once or twice.
  • A. Piece Selection: Selection also required scores to permit tempo rubato, dynamic variation, ornamentation, and other self-interpretations.This criterion was intended to avoid rigid performances.

B. Recruiting Musicians

URMP recruited a conductor, pianist, and 22 instrumental players from University of Rochester-affiliated music communities. Conducting videos provided a shared expressive synchronization basis for separately recorded instrumental parts.

  • Musician recruitment: 22 instrumental players recorded all instrumental parts, drawn from Eastman School of Music students and University of Rochester ensembles and orchestras.The conductor had more than 20 years of conducting experience, and the pianist was a graduate student in piano performance.
  • Conducting videos: Each piece began with a conducting video showing a conductor and pianist performing together on a Yamaha grand piano.These videos served as the basis for synchronizing the different instrumental parts.
  • Expressive reference: The conductor and pianist set each piece’s tempo after rehearsing, retained repeats, and implemented all score expression markings.Players later followed the conducting video, preserving expressive performance choices without requiring full-ensemble rehearsals.
  • Recording procedure: Players followed the conducting video on a laptop and listened through a low-latency Bluetooth earphone; difficult pieces could require several recording shots.Simpler pieces were completed in one shot before quality approval.

E. Mixing and Assembling Individual Recordings

The assembly workflow aligned independently recorded audio and video, manually synchronized instrumental tracks, balanced their loudness, and composited performers into a concert-hall setting while producing pitch annotations.

  • Audio-video alignment: Camera audio was replaced with stand-alone high-quality audio, then the two recordings were automatically aligned using Final Cut Pro’s “synchronize clips” function.Independent camera and microphone control created a relative shift that required correction.
  • Track assembly: Individual instrumental recordings were manually time-shifted against one another by focusing on fast sections with clear note onsets.Some tracks also received manual loudness adjustments to improve volume balance.
  • Video compositing: Color correction, spatial masks, and chroma keying separated players from the unevenly lit blue background before video compositing.The corrected foregrounds were used to reduce artifacts caused by shadows and background texture variation.
  • Ground-truth annotation: Pitch annotations were generated separately for each audio track using Tony, which implements pYIN for frame-wise monophonic F0 estimation.Pitch trajectories used a 5.8 ms frame hop before interpolation to 10 ms, and note sequences were extracted with HMM Viterbi decoding.

V. THE DATASET

URMP organizes 44 pieces with synchronized multimodal recordings, scores, and track-level annotations, supporting both conventional MIR and audio-visual analysis. The dataset’s synchronization is evaluated against Bach10 and WWQ.

  • Dataset contents: Each of the 44 pieces is organized with a MIDI score, PDF sheet music, individual and mixed audio, assembled video, and ground-truth annotations.The complete dataset is deposited as a 12.5 GB collection in the Dryad Digital Repository.
  • Audio and score: Individual and mixed WAV recordings use 48 KHz sampling and 24-bit depth, with track naming aligned to score order.This organization connects each isolated source to its corresponding score track.
  • Video: Assembled MP4 videos use H264 encoding at 1080P resolution and 29.97 FPS, with players rendered left-to-right in score-track order.The assembled videos preserve an explicit spatial ordering of performers.
  • Annotations: Annotations provide ground-truth frame-level pitch trajectories and note-level transcriptions for individual tracks.These annotations are supplied in ASCII-delimited text format.
  • Synchronization comparison: URMP synchronization is compared with Bach10 and WWQ using onset deviations for score-notated simultaneous notes.PHENICX-Anechoic is excluded because it used the same approach as URMP and contains much larger symphony ensembles.

1) Quantitative Evaluation:

URMP’s synchronization was evaluated numerically and subjectively, while its videos were assessed for occlusion and ROI resolution. Results place URMP between WWQ and Bach10 on synchronization and quantify usable visual detail.

  • Quantitative synchronization: 20 to 60 ms is URMP’s maximum onset-deviation range, compared with 60 to 80 ms for Bach10; WWQ ranks best overall.The evaluation uses the maximum deviation among score-notated simultaneous notes when polyphony exceeds two.
  • Subjective synchronization: 9 of 32 subjective rankings placed URMP first, while 17 placed it second, consistent with the quantitative comparison.Eight subjects each provided four rankings after listening to randomly selected triplets.
  • Evaluation caveat: Onset-time evaluation is limited by ambiguity in identifying onset instances for some soft articulations.The subjective evaluation supplements this limitation with listener-based synchronization rankings.
  • Spatial occlusion: The fixed right-side camera view leaves right-side faces unobstructed, but hand and arm occlusion varies by instrument type.Cello and bass left arms are sometimes occluded by the instrument body, while brass fingering hands remain visible.
  • ROI resolution: About 100×100, 70×70, and 40×40 pixels are available for faces, hands, and mouths, respectively, in assembled 1080P video.Resolution decreases slightly for quintets compared with duets because more players share the frame.

D. Limitations of the Dataset

URMP’s main limitations concern its visual realism and coverage: single-camera recordings restrict viewpoints, omit some objects and interactions, and introduce assembly artifacts.

  • Visual limitations: Single-camera videos limit viewpoint diversity, sometimes exclude important objects, and make player arrangements less natural than real chamber performances.The fixed view can hinder pitch inference from violin fingering and place bow endpoints or players’ heads outside the frame.
  • Assembly limitations: Unmodeled occlusions may make video analysis easier than in real scenarios.Occlusions between players or by music stands were not considered during assembly.
  • Assembly limitations: Because instrumental parts were recorded in isolation, the assembled videos lack natural player interactions and cannot support visual analysis of those interactions.Eye contact and body-motion interactions commonly observed in real performances are absent.
  • Production artifacts: Minor production artifacts include irrelevant movements from the earphone wire and slight foreground color changes from chroma keying.These issues are identified as avoidable in future dataset creation.
  • Supported applications: URMP nevertheless supports a broad range of audio-only and audio-visual MIR tasks, including multi-pitch analysis and score-informed source separation.The dataset section highlights existing audio tasks and novel tasks requiring both modalities.

1) Multi-pitch Analysis:

The benchmark evaluates multi-pitch analysis and score-informed source separation on URMP against Bach10, while also motivating audio-visual extensions. URMP is harder for pitch analysis, but score information yields similar separation performance for matched conditions.

  • Multi-pitch analysis: TP, FP, and FN define accuracy evaluation by comparing estimated and ground-truth pitches within a quarter-tone tolerance.TP denotes true positives, FP false positives, and FN false negatives.
  • Multi-pitch analysis: MPE and MPS accuracies decrease as URMP polyphony increases, with quartet results significantly below Bach10.URMP’s broader musical variety and repeated instruments make timbre-based pitch streaming more difficult.
  • Score-informed source separation: Score-informed separation performance is very similar on URMP and Bach10 for quartets with the same polyphony and instrument tracks.The reported result indicates that score information helps address URMP’s greater challenges relative to Bach10.
  • Audio-visual tasks: The dataset also supports new MIR tasks requiring both audio and visual modalities, with evaluation strategies and baseline systems provided.These tasks are presented as directions for further research.

1) Visually Informed Multi-pitch Analysis:

The dataset enables visually informed multi-pitch and polyphonic vibrato analysis by linking performance motion with audio and score information. Baseline results show visual cues can improve pitch analysis and support robust vibrato detection and parameter estimation under polyphony.

  • Visually Informed Multi-pitch Analysis: Visual information can significantly help multi-pitch analysis by revealing fingering and player activity that inform note prediction and source assignment.
  • Visually Informed Multi-pitch Analysis: The baseline models each string player’s play/non-play activity from video and uses it to constrain audio-based pitch analysis.
  • Visually Informed Multi-pitch Analysis: 2-12% improvement is observed across multi-pitch tasks and pieces compared with the audio-based method.
  • Polyphonic Vibrato Analysis: Visual vibrato cues remain useful under polyphony because string-player finger motion directly reflects pitch fluctuation without degrading as sources increase.
  • Polyphonic Vibrato Analysis: The vibrato task covers 19 pieces and estimates vibrato-note labels plus rate and extent from pitch-contour autocorrelation annotations.
  • Polyphonic Vibrato Analysis: The video baseline generally exceeds 90% F-measure for vibrato-note detection, with average errors of 0.38 Hz for rate and 3.47 musical cents for extent.90% of errors fall within 1 Hz and 10 musical cents, respectively.
  • Polyphonic Vibrato Analysis: The current vibrato task is limited to vibrato analysis, while broader playing-technique detection and extension to non-string instruments remain anticipated directions.

3) Other Emerging New Tasks:

URMP supports emerging audio-visual tasks beyond the paper’s benchmarked analyses. These directions use learned relations between sound events and visual movements or objects to separate, associate, or generate modalities.

  • Other Emerging New Tasks: Visually informed source separation can leverage associations between audio events, such as violin notes, and visual movements, such as bowing.
  • Other Emerging New Tasks: Audio-visual source association links sound sources or notes to visual objects such as players and could support sound-track targeting from visual scenes.
  • Other Emerging New Tasks: Audio-visual cross-modality generation seeks to generate one modality from the other by modeling their relations, with temporal dependencies identified as an open direction.
  • Other Emerging New Tasks: URMP provides a multimodal performance resource for source separation, transcription, audio-score alignment, performance analysis, and related research applications.
Loading 1612.08727v3…