Source-linked AI summary
A Framework for the Robust Evaluation of Sound Event Detection
Cagdas Bilen, Giacomo Ferroni, Francesco Tuveri, Juan Azcarreta, Sacha Krstulovic
TL;DR
Conventional SED evaluation is limited by operating-point dependence and collar-based decisions that are sensitive to subjective event boundaries. The paper introduces robust TP/FP definitions, PSD-ROC curves, and PSDS, and demonstrates broader system-performance insight and flexible evaluation across criteria.
Problem
Conventional SED metrics remain limited by operating-point dependence and collar-based timing decisions that are sensitive to labelling subjectivity.
Method
The framework redefines true and false positives and uses PSD-ROC curves plus PSDS to evaluate polyphonic SED across operating points and user-experience criteria.
Results
The framework provides more robust measurements against labelling subjectivity and more comprehensive system-performance insight than single-operating-point F1 scores.
Takeaways & Limitations
PSDS rankings and score differences can vary with evaluation criteria, enabling comparisons tailored to application and user-experience requirements.
Takeaways & Limitations
Collar-based evaluation can force reasonable alternative interpretations of temporal event structure to count as classification errors.
Abstract
from arXiv · showhide
This work defines a new framework for performance evaluation of polyphonic sound event detection (SED) systems, which overcomes the limitations of the conventional collar-based event decisions, event F-scores and event error rates. The proposed framework introduces a definition of event detection that is more robust against labelling subjectivity. It also resorts to polyphonic receiver operating characteristic (ROC) curves to deliver more global insight into system performance than F1-scores, and proposes a reduction of these curves into a single polyphonic sound detection score (PSDS), which allows system comparison independently from operating points (OPs). The presented method also delivers better insight into data biases and classification stability across sound classes. Furthermore, it can be tuned to varying applications in order to match a variety of user experience requirements. The benefits of the proposed approach are demonstrated by re-evaluating the baseline and two of the top-performing systems from DCASE 2019 Task 4.
1. INTRODUCTION
Sound event detection evaluates automatically detected sounds in challenging, polyphonic audio streams, but conventional metrics remain sensitive to operating points, subjective event boundaries, and dataset biases. The paper proposes more robust TP/FP definitions, PSD-ROC curves, and PSDS for broader system evaluation.
- Sound event detection automatically detects sound events from audio streams, supporting applications including smart homes, smart speakers, headphones, and mobile devices.
- Conventional event-wise and segment-wise metrics still depend on operating-point tuning, which can change system rankings under the same metric.ROC, DET, and AUC methods evaluate systems globally across operating points, but this practice is less established in SED.
- Collar-based event metrics emphasize start and end times even though human labellers may reasonably interpret temporal event structure differently.For example, repeated dog barks may be labelled as one event or several separate events.
- Cross-triggers are false positives matching another labelled class, and distinguishing them from raw false positives can reveal biases affecting acoustically similar classes.This distinction helps assess whether false positives reflect data bias rather than an acoustic modelling defect.
- The paper proposes robust TP and FP definitions, PSD-ROC curves, and PSDS to evaluate SED systems more globally than single operating-point metrics.The proposed framework is designed to address operating-point variation, labelling subjectivity, and possible evaluation-dataset biases.
2. BACKGROUND
Polyphonic SED evaluates simultaneous events from multiple sound classes by comparing system detections with ground-truth labels. The paper formalizes this task and describes operating points and collar-based matching, whose boundary dependence can make evaluation vulnerable to reasonable labelling variation.
- 2.1. Definition of the Sound Event Detection task: Polyphonic SED detects events from multiple classes while allowing events to occur simultaneously.
- 2.1. Definition of the Sound Event Detection task: The class set C contains sound classes, while each ground-truth label and detection is associated with a class and temporal start and end times.
- 2.1. Definition of the Sound Event Detection task: The evaluation task measures the performance of a system that outputs operating-point-dependent detection sets given ground-truth event labels.Ground-truth labels and detections are represented by class, start time, and end time.
- 2.1. Definition of the Sound Event Detection task: Operating-point parameters τc adjust system permissiveness; higher thresholds generally produce more conservative detections, while lower thresholds increase permissiveness.
- 2.2. Limitations of conventional collar-based evaluation: Conventional collar evaluation counts a detection as a true positive when its timing satisfies predefined collar constraints relative to a ground-truth event.The collar duration may be fixed or proportional to ground-truth duration, and an end-time collar may be omitted when end times are difficult or irrelevant to label.
- 2.2. Limitations of conventional collar-based evaluation: Collars can force reasonable alternative interpretations of an event’s temporal structure to produce classification errors.Repeated dog barks may reasonably be treated as one event or as several separate events.
3. PROPOSED EVALUATION FRAMEWORK
The framework replaces collar-based event decisions with tolerance criteria and separates cross-triggers from other false positives. It then builds operating-point-independent polyphonic ROC curves and summarizes them with PSDS, while allowing application-specific costs and stability preferences.
- Robust event decisions: The framework redefines true positives and false positives using Detection Tolerance Criterion and Ground Truth intersection Criterion rather than boundary collars.DTC filters relevant detections using an intersection tolerance, while GTC identifies correctly detected ground-truth events using a ground-truth tolerance.
- Robust event decisions: Intersection tolerances are intended to improve robustness to disagreements at sound-event boundaries, where human labellers may interpret temporal structure differently.The approach is motivated by disagreements involving fades, repeated units, and other boundary interpretations.
- Robust event decisions: Cross-Trigger Tolerance Criterion counts cross-triggers separately from false positives to expose confusions between target sound classes and possible dataset biases.Cross-triggers are false positives that match another labelled class, and CTTC introduces a separate count for them.
- Metrics relevant to user experience: The framework measures TP performance as a detected-event proportion, while false-positive and cross-trigger performance are expressed as rates per unit time.The effective false-positive rate can weight cross-triggers according to their user-experience cost through αCT.
- ROC construction and PSDS: For each class, dominated operating points are discarded, remaining points are interpolated into ROC curves, and class-dependent curves are averaged into one polyphonic ROC curve.This removes dependence on the operating point when constructing the curve and retains best-case trade-offs between TP ratio and effective false-positive rate.
- ROC construction and PSDS: The effective TP ratio combines the mean and standard deviation of class-wise TP ratios, and the normalized area under the resulting PSD ROC curve defines PSDS.αST controls the cost assigned to instability across classes, while emax sets the maximum effective false-positive rate of interest.
4. EXPERIMENTAL RESULTS
The experiments re-evaluate three DCASE 2019 Task 4 systems using collar-based and DTC/GTC F1-scores, PSD-ROC curves, and PSDS under varied settings. The results show that TP/FP definitions and evaluation criteria can change rankings and reveal operating-point, cross-trigger, and class-stability differences.
- Experimental setup: Three publicly available DCASE 2019 Task 4 systems were evaluated with both the collar-based metric and the proposed PSDS metric.The systems were the baseline, first-ranking, and fourth-ranking systems, renamed Systems 1, 2, and 3.
- Robustness of a DTC/GTC-based F1-score: Reinterpreting TPs and FPs with DTC/GTC changes system rankings and makes the systems perform closely under the corresponding F1-score.Increasing the DTC/GTC tolerance parameters to 0.8 further evens out the performance measures.
- Summarising performance into a single PSDS figure: PSDS evaluates performance across operating points and incorporates false-positive rates per unit of time, providing broader insight than single-operating-point F1-scores.An open-source implementation of PSDS is available.
- Summarising performance into a single PSDS figure: System 2 performs better than Systems 1 and 3 across all PSD-ROC performance trade-offs, while System 1 is preferable at very low FP rates and System 3 at higher TP ratios.The comparison uses varied thresholds to create PSD-ROC curves, with cross-triggers and class stability ignored in this setting.
- A flexible evaluation criterion: Average scores, score differences, and system rankings change under different evaluation criteria, reflecting application-dependent user-experience requirements.The results also indicate unstable classification performance across classes when αST = 1, while cross-trigger weighting can favor System 1 over System 3 in specified settings.
5. CONCLUSIONS
The paper presents an SED evaluation framework designed to address operating-point variation, labelling subjectivity, and evaluation-dataset biases. Its TP/FP definitions, PSD-ROC curves, and PSDS provide more robust and comprehensive evaluation while allowing application-specific criteria.
- Contributions: The framework redefines TPs and FPs using intersection criteria rather than event boundaries, improving robustness to labelling subjectivity.Its benefits are demonstrated by re-evaluating the DCASE 2019 Task 4 baseline and two top-ranking systems on the same task and data.
- Contributions: PSD-ROC curves and PSDS assess performance across operating-point ranges and incorporate false-positive rates per unit of time, providing more comprehensive insight than F1-scores.The evaluation parameters can be adjusted to match different applications and user-experience requirements.