Source-linked AI summary
FSD50K: An Open Dataset of Human-Labeled Sound Events
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, Xavier Serra
TL;DR
SER research lacks sufficiently open, broad, and reusable dataset resources for benchmarking and model development. FSD50K combines Freesound audio, the AudioSet Ontology, manual labeling, and careful evaluation-set curation into an open benchmark. It contains 51,197 clips across 200 classes, while retaining label-noise and real-world-naturalness limitations.
Problem
SER research lacks sufficiently open, broad, and reusable dataset resources for benchmarking and model development.
Method
FSD50K combines Freesound audio, the AudioSet Ontology, manual labeling, and careful evaluation-set curation into an open benchmark.
Results
FSD50K contains 51,197 clips totaling over 100h, manually labeled across 200 AudioSet Ontology classes.
Takeaways & Limitations
Creative Commons licensing makes FSD50K freely distributable, including audio waveforms, for SER research and benchmarking.
Takeaways & Limitations
FSD50K retains missing “Present” labels, especially in development data, because labels are validated from automatically nominated candidates.
Abstract
from arXiv · showhide
Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2M tracks from YouTube videos and encompassing over 500 sound classes. However, AudioSet is not an open dataset as its official release consists of pre-computed audio features. Downloading the original audio tracks can be problematic due to YouTube videos gradually disappearing and usage rights issues. To provide an alternative benchmark dataset and thus foster SER research, we introduce FSD50K, an open dataset containing over 51k audio clips totalling over 100h of audio manually labeled using 200 classes drawn from the AudioSet Ontology. The audio clips are licensed under Creative Commons licenses, making the dataset freely distributable (including waveforms). We provide a detailed description of the FSD50K creation process, tailored to the particularities of Freesound data, including challenges encountered and solutions adopted. We include a comprehensive dataset characterization along with discussion of limitations and key factors to allow its audio-informed usage. Finally, we conduct sound event classification experiments to provide baseline systems as well as insight on the main factors to consider when splitting Freesound audio data for SER. Our goal is to develop a dataset to be widely adopted by the community as a new open benchmark for SER research.
I. INTRODUCTION
SER research needs larger, broader, and openly reusable datasets than many existing resources provide. FSD50K addresses this gap with a human-labeled, freely distributable benchmark built from Freesound.
- SER identifies sound classes in audio and supports applications including healthcare, urban planning, bioacoustics, surveillance, and noise monitoring.
- Larger and more comprehensive datasets are needed to support data-hungry deep-learning approaches, whereas earlier SER datasets had limited size and coverage.
- Recent post-AudioSet datasets remain task- or domain-specific, have limited class coverage, or use synthetic audio.
- 51,197 clips totaling over 100h are manually labeled across 200 AudioSet Ontology classes in FSD50K, with Creative Commons licensing enabling waveform distribution.
- FSD50K contributes an open dataset, documented creation process, dataset characterization, and classification baselines addressing data-informed SER design and splitting decisions.
- Unlike AudioSet’s feature-based official release, FSD50K provides freely distributable audio waveforms, while retaining a broad sound-event scope.
C. Datasets Released After AudioSet
FSD50K is positioned as an open, expandable alternative to task-specific or domain-specific SER datasets. Its design prioritizes label quality and evaluation reliability while acknowledging the trade-offs of weak labeling and manual annotation.
- C. Datasets Released After AudioSet: Post-AudioSet SET datasets often target particular tasks or domains, motivating an open dataset with broader coverage and expansion potential.
- C. Datasets Released After AudioSet: FSD50K is designed to be fully distributable, cover many everyday sounds, and expand in both data and vocabulary.
- C. Datasets Released After AudioSet: Weak clip-level labels are simpler and less time-consuming to gather than onset/offset labels, but they impose limitations on training and evaluation.
- 3) Emphasis on Evaluation Set:: The authors prioritize a comprehensive, diverse, reliably annotated, and representative evaluation set because it defines target behavior for benchmarking.
- 3) Emphasis on Evaluation Set:: Manual annotation is chosen over active learning to favor reliable labels and control over labels and sound predominance, despite greater human effort.
B. Overall procedure
FSD50K construction combines Freesound audio, the AudioSet Ontology, and Freesound Annotator through automated candidate mining followed by progressive filtering and annotation preparation.
- B. Overall procedure: The creation pipeline starts from Freesound and the AudioSet Ontology and progressively filters audio clips and vocabulary classes before producing FSD50K.
- B. Overall procedure: Freesound supplies diverse audio clips and user metadata, while the AudioSet Ontology supplies a broad hierarchical vocabulary and Freesound Annotator supports dataset management and annotation.
- B. Overall procedure: Candidate clips are automatically retrieved by matching Freesound user tags to class-specific keywords derived from ontology descriptions and frequent co-occurring tags.
- B. Overall procedure: 268,261 clips remained after filtering out clips longer than 90s, with an average of 2.62 candidate labels per clip.
- B. Overall procedure: The automatically nominated labels indicate potential sound-event presence and require manual validation because tag choices and class ambiguity can induce errors.
E. Validation Task
The validation task uses a two-phase human annotation workflow supported by training materials, prioritized clip presentation, and quality-control mechanisms to refine automatically nominated labels.
- E. Validation Task: Raters first learn each class through its ontology location, textual description, and representative examples, then validate candidate clips for that class.
- E. Validation Task: The final tool was redesigned after an internal quality assessment involving 11 subjects who validated 12 candidates per class and provided feedback.
- E. Validation Task: Class FAQs help resolve ambiguous descriptions, while spectrograms and loudness normalization support more consistent and less burdensome annotation.
- E. Validation Task: Separating “Present and predominant” from “Present but not predominant” distinguishes isolated clean events from clips with multiple or acoustically adverse events.
- E. Validation Task: Verification clips trigger response discard after errors, and candidate labels require agreement between two different raters before becoming ground truth.
- E. Validation Task: The system prioritizes unresolved label-clip pairs and short clips, whose higher label density makes them more informative for learning.
4) Annotation Campaign:
The annotation campaign combined crowdsourcing and hired raters, with strategies adapted to class difficulty and procedures designed to improve consistency. It produced a large pool of mostly correct labels, while retaining possible incompleteness from tag-based candidate nomination.
- Classes were divided by estimated difficulty, with annotation strategies assigned separately to easy, medium, and harder subsets.The campaign used both crowdsourcing and hired raters because some classes were more difficult to annotate than others.
- Over 350 raters contributed, including voluntary participants, six hired raters, and the paper’s first three authors.
- Candidate nomination based on Freesound tags produced mostly correct labels but could miss sound events absent from user-generated tags.The labels were estimated at 94.3% correctness in the supplied passage, while incompleteness remained possible.
- The split prioritized uploader-disjoint development and evaluation sets, allocating small-uploaders’ content to evaluation for greater diversity.The rationale included diversity of sound sources, acoustic environments, and recording equipment.
- The evaluation allocation favored longer recordings because they were expected to contain more events and be more representative of real-world audio.
2) Split Method:
The split method constructs uploader-disjoint candidate development and evaluation pools through iterative, score-based allocation. It prioritizes diversity and class coverage rather than exact fine-level class matching.
- The method iteratively allocates uploaders’ content to evaluation after ranking uploaders by an ad hoc score.Off-the-shelf random sampling, iterative stratification, and combinatorial optimization were considered unsuitable under the stated constraints.
- Uploader scores combine the maximum class-specific label count with the uploader’s average labels per touched class.The score discourages concentration in one class and favors uploaders contributing more diversely across classes.
- The algorithm traverses 113 leaf classes, beginning with the least represented, and allocates ranked uploader content until each target is reached.Targets are proportional to class label counts and constrained to 50–100 labels per class.
- Each uploader is limited by default to contributing 0.1t_ci clips to class c_i in the evaluation set.
- The resulting candidate development and evaluation pools are disjoint in terms of uploaders, and the candidate evaluation pool is exhaustively labeled afterward.
G. Refinement Task
The refinement task exhaustively reviews candidate evaluation labels and adds missing labels by enabling ontology-based exploration. Its output is an exhaustively labeled pool that supports subsequent curation into the released dataset.
- Refinement addresses potentially incomplete evaluation labels that could penalize classifiers for predicting correct but missing events.
- The annotation tool lets raters review existing labels and add missing “Present” labels.
- Clips are grouped by sound class, and raters explore the ontology to assign the most specific labels possible, typically leaf labels.
- Raters receive training before refinement, while the validation task’s quality-assurance practices are also applied.
- Each evaluation label is verified or reviewed by between two and five annotators across validation and refinement.
- The stage produces exhaustively labeled clips for the considered vocabulary, with most forming the evaluation set.
- Post-processing merges low-prior classes into parent classes, yielding the 200-class ground-truth vocabulary released in FSD50K.The raw sound-collection format retains all generated labels, including classes with little data.
IV. DATASET DESCRIPTION
FSD50K contains 51,197 clips across 200 AudioSet Ontology classes, with separate uploader-independent development and evaluation splits. Its variable-length, weakly labeled audio and differing annotation completeness shape how the dataset should be interpreted.
- Dataset composition: FSD50K contains 51,197 clips distributed across 200 AudioSet Ontology classes.The classes include 144 leaf nodes and 56 intermediate nodes, with unequal class distributions.
- Dataset distributions: Eval clips tend to have more labels and last slightly longer than dev clips.The label-distribution comparison uses unpropagated labels, while clip-length distributions use 1/3-second bins.
- Dataset splits: The dev and eval splits contain no clips from the same uploader.This separation is intended to prevent uploader overlap between development and evaluation data.
- Annotation completeness: Eval annotations are exhaustive, whereas most dev annotations are correct but potentially incomplete.The reported label counts use unpropagated labels, which exclude propagated ancestors and can slightly undercount human-provided labels in some cases.
- Labels and clip length: FSD50K uses variable-length clips from 0.3 to 30 seconds with clip-level weak labels.The dev set is roughly 70% weakly labeled and 30% strongly labeled, while longer clips increase uncertainty about event location.
- Audio quality: FSD50K’s heterogeneous Freesound recordings make objective audio-quality claims difficult, although its mean SNR exceeds AudioSet’s.The paper treats P.563-derived SNR as only a rough indication because P.563 was designed for human speech.
3) Real-world Audio:
FSD50K combines diverse Freesound recordings with substantial annotation and class-imbalance constraints. The authors identify missing labels, recording-condition mismatch, uploader-related bias, and merged classes as important boundaries for use and evaluation.
- 3) Real-world Audio: Freesound recordings range from real-world events to carefully staged or purposely generated sounds, which may mismatch deployment conditions.The effect of this potential mismatch on generalization to adverse scenarios remains an open question.
- 1) Label Noise: Missing “Present” labels affect dev more because its validation process depends on previously nominated labels.Events not nominated by the system can remain unlabeled, producing false negatives.
- 1) Label Noise: 50.9% of 11,847 processed clips received at least one additional label, while 5.7% of incoming labels were rejected.The refinement results indicate substantial missing-label noise alongside 94.3% verified correctness for incoming labels.
- 2) Class Imbalance: Class imbalance arises from uneven class abundance, variable clip lengths, and the ontology hierarchy.Data scarcity and low nomination performance also leave some classes much less represented.
- 4) Data Bias: A few classes may contain development data dominated by large uploaders, creating a potential learnable data bias.The paper highlights musical-instrument classes such as Trumpet as an example.
- 5) Ontology Constraints: Some ontology leaf nodes were merged with parent classes because of data scarcity.Examples include Blender, Chopping (food), and Toothbrush merged into Domestic sounds, home sounds.
- Dataset Uses: FSD50K supports multilabel large-vocabulary classification and several sound-recognition tasks and challenges.The dataset has supported DCASE tasks involving noisy labels, minimal supervision, weak labels, and sound separation.
E. FSD50K and AudioSet
FSD50K and AudioSet share the AudioSet Ontology but differ in openness, waveform availability, labeling, and evaluation annotation completeness. The section also introduces metrics and experiments for assessing multi-label sound event tagging.
- FSD50K uses a smaller AudioSet Ontology subset and provides waveform audio under several Creative Commons licenses, whereas AudioSet officially releases pre-computed features.
- FSD50K adds event-predominance annotations and exhaustive evaluation labels, while AudioSet primarily provides presence annotations and generally incomplete evaluation labels.
- FSD50K and AudioSet are highly heterogeneous, but Freesound clips are typically recorded to capture audio, whereas YouTube recordings may have greater device and real-world-condition diversity.
- The experiments evaluate multi-label sound event tagging with a baseline pipeline and examine challenges in splitting Freesound audio for recognition tasks.
- The evaluation uses within-class mAP and d′ together with between-class label-weighted label-ranking average precision, with larger values indicating better performance.
B. Baseline Systems
The baseline systems compare common neural architectures on FSD50K using log-mel spectrogram patches and clip-level aggregation. VGG-like achieves the strongest overall performance, while CRNN leads label-ranking performance and differs by sound class.
- Learning Pipeline: Incoming audio is downsampled to 22.050 kHz and converted into 96-band log-mel spectrogram patches of 1 second and shape 101x96.
- Learning Pipeline: Clip-level predictions average per-class scores across time-frequency patches, and validation uses the same aggregation because patch-level inherited labels were misleading.
- Architectures: All models end with 200 sigmoid outputs for multi-label classification, while CRNN and VGG-like designs receive limited preliminary tuning.
- Results: VGG-like is best overall across metrics despite being lighter, DenseNet-121 follows, CRNN has the best lωlrap, and ResNet-18 performs worst.
- Results: CRNN generally has lower per-class AP than VGG-like but performs better for several speech, human-vocalization, and animal-vocalization classes with marked temporal behavior.
C. Impact of Train/Validation Separation
The study compares three validation splits and shows that uploader contamination can separate validation from evaluation performance. The proposed split minimizes within-class contamination and is preferred for benchmarking.
- The comparison trains and evaluates CRNN models using val random, val is, and the proposed val splits.
- Val random and val is are constructed through repeated candidate separations, using distribution similarity or shared-uploader minimization respectively.
- Val random and val is contain within-class and between-class contamination, whereas the proposed val minimizes within-class contamination and retains mostly between-class contamination.
- Validation performance is substantially better than evaluation performance for val random and val is when training and validation share uploaders and classes.
- The proposed val split has slightly higher train performance and slightly lower validation and evaluation performance, but is selected for benchmarking because contamination is minimized.
- Benchmarking should explicitly specify the validation split because Freesound data separation can non-negligibly affect learning and performance.
APPENDIX A ONTOLOGICAL NOMENCLATURE
The appendix defines the ontology’s hierarchical terminology and describes how FSD50K reduces the candidate vocabulary to 200 semantically usable classes. The released dataset retains more specific annotations in its collection format.
- APPENDIX A ONTOLOGICAL NOMENCLATURE: The AudioSet Ontology’s 632 classes are called nodes, divided into leaf nodes at the hierarchy bottom and intermediate nodes with children.
- Determine FSD50K vocabulary: The refined data contain a candidate development set with potentially incomplete labels and an exhaustively labeled candidate evaluation set across 395 classes before post-processing.
- Determine FSD50K vocabulary: Valid leaf nodes require at least 100 clips and no extreme development/evaluation imbalance, balancing per-class data abundance against vocabulary breadth.
- Determine FSD50K vocabulary: Non-valid leaf nodes are merged with parents, with the merge strategy depending on whether valid siblings remain in the hierarchy.
- Determine FSD50K vocabulary: Some valid leaves are removed to create a more semantically consistent vocabulary, while more specific annotations remain available in the sound collection format.
- Determine FSD50K vocabulary: Discarding abstract or ambiguous intermediate nodes yields 200 classes comprising 144 leaf nodes and 56 intermediate nodes.
C. Validation Set
The proposed validation split balances class stratification, contamination control, and target proportions, but these criteria cannot all be satisfied strictly for Freesound data. The resulting candidate set uses uploader-based allocation and explicitly trades off stratification against contamination.
- Design criteria: A useful validation set should preserve class stratification, limit contamination across splits, and target roughly 10–20% of development data.For multilabel, variable-length Freesound audio, the proportion may differ when measured by clips, labels, or duration.
- Design criteria: Random sampling and iterative stratification can match proportions and class distributions but may split uploader content across train and validation.Because uploader sizes vary widely, preserving uploader non-divisibility can instead distort the target class distribution.
- Allocation procedure: The allocation algorithm ranks uploaders per class and transfers low-score uploaders first until each class reaches its validation target.The score combines the uploader’s class-specific data amount with its average labels per class, weighted by tunable α and β parameters.
- Allocation procedure: Within-class contamination can arise when uploader content is split to reach a class target, while between-class contamination is not addressed unless multilabel clips connect the classes.The procedure permits within-class contamination when adding an uploader would overshoot the target and the current amount is at most 0.75t_ci.
- Outcome: 4170 clips form the candidate validation set, representing 10.2% of the development set and a tradeoff between stratification and contamination.Among 2224 uploaders represented in validation, 641 also contribute content to training, mostly through between-class contamination.
D. Hierarchical Label Propagation
FSD50K propagates labels upward through the AudioSet ontology to create hierarchical annotations, but ambiguous parent classes prevent complete paths in some clips. The dataset therefore omits uncertain ancestors and partially reviews critical evaluation cases.
- Propagation: Current lower-level labels are propagated upward by identifying ancestors along each class’s hierarchical path.This produces exhaustive hierarchy-wise labeling across train, validation, and evaluation sets when parent information is available.
- Ambiguity: Missing disambiguating parent information makes some hierarchical paths incomplete, such as Growling connecting directly to Animal when the source animal is unknown.Raters were instructed to provide disambiguating parents in refinement tasks, but they did not always specify them.
- Ambiguity: The policy omits ambiguous ancestors rather than risk incorrect labels, leaving development cases as-is while partially reviewing and correcting critical evaluation cases.Filtering beyond the 200 selected classes also means some clips retain only part of their ancestor path or a single label.