Source-linked AI summary
Towards Actionable Surgical Team Dynamics: from Teamwork to Counterfactual Annotations
Vincenzo Marco De Luca, Antonio Longa, Andrea Passerini
TL;DR
Existing operating-room datasets provide limited integrated coverage of the multimodal signals and teamwork constructs needed to study collaboration in high-stakes settings. The paper extends a public OR dataset with speaker-aware and language-aware representations, multilevel teamwork annotations, and counterfactual annotations. The resulting resource supports integrated, multimodal, and explainable modeling of surgical teamwork, while its counterfactual layer remains subjective rather than experimentally causal.
Problem
Existing datasets provide limited integrated coverage of multimodal teamwork signals, while teamwork involves dynamic communication, coordination, leadership, situational awareness, and role-dependent behavior.
Method
The paper extends a public operating-room dataset with enriched representations, multilevel annotations, and a counterfactual annotation layer for teamwork analysis.
Results
The resulting resource supports integrated, multimodal, and explainable modeling of teamwork through speaker-aware preprocessing, bilingual transcripts, teamwork annotations, and counterfactuals.
Takeaways & Limitations
The dataset supports research on temporal-relational learning, multi-task modeling, and human-centered AI for high-stakes collaborative settings.
Takeaways & Limitations
The enrichment and annotation process relies heavily on expert and manual effort, limiting scalability, and counterfactual annotations remain subjective rather than experimentally validated causal mechanisms.
Abstract
from arXiv · showhide
Modeling team interactions in high-stakes environments such as operating rooms is critical for understanding how coordination, communication, and individual behaviors shape team performance and safety outcomes. Existing datasets in this domain are often fragmented across modalities, annotation schemes, and formats, limiting their ability to support integrated analyses of real-world collaborative processes. We address this limitation by introducing an extended multimodal dataset for surgical team interaction analysis, built from real operating room recordings. Starting from an existing corpus, we construct an analysis-ready version of the data by providing speaker diarization, transcripts, and multi-level annotations capturing team performance, interaction processes, and individual characteristics. Team performance is assessed using a standardized surgical teamwork evaluation protocol, while interaction quality and individual attributes are annotated through structured rating schemes covering collaboration, group dynamics, and non-technical skills. To further support the study of coordination breakdowns and performance variability, we introduce counterfactual annotations that describe plausible alternative team outcomes in the presence of observed interaction failures, enabling analysis of how specific behavioral patterns may relate to different trajectories of team performance. In addition, we provide structured temporal and relational representations designed to support computational modeling of teamwork processes and the design of AI-assisted collaborative systems. The dataset is designed to support the study of how individual actions, interaction patterns, and team-level processes jointly contribute to team outcomes in surgical settings, providing a unified resource for analyzing collaborative behavior in high-stakes domains.
1 Introduction
Teamwork is essential to safety and performance in operating rooms, yet existing surgical datasets provide limited multimodal and teamwork-oriented coverage. This paper introduces a benchmark extending such data with richer representations, multilevel annotations, and counterfactuals.
- Teamwork in operating rooms depends on communication, coordination, leadership, situational awareness, and mutual support alongside technical proficiency.
- Public surgical datasets remain limited or restricted in modality, although teamwork unfolds through speech, turn-taking, silence, interruptions, body conduct, roles, and extended coordination patterns.
- The proposed benchmark builds on a public operating-room dataset and adds speaker-aware and language-aware information for surgical teamwork analysis.
- Its annotation framework covers team-, interaction-, and individual-level assessment, while counterfactual annotations identify behaviors associated with teamwork degradation.
- The paper organizes the resource construction, annotation protocol, benchmark tasks, baselines, and limitations into a complete analysis pipeline.
2 Related Work
Prior datasets and teamwork frameworks offer complementary but disconnected views of clinical collaboration. The paper addresses this gap with richer multimodal representations and annotations spanning team, individual, interaction, and counterfactual levels.
- Existing clinical team-modeling resources and performance frameworks remain largely disconnected, limiting comprehensive modeling of team dynamics.
- Common surgical benchmarks emphasize visual perception and workflow tasks rather than communication, coordination, or role dynamics.
- PoPCaP lacks real operating-room video and public availability, while CliniDial excludes video and therefore restricts analysis to verbal interactions.
- The proposed extension extracts multimodal features for operating-room scenarios and team-modeling settings through richer data representations.
- Structured teamwork frameworks such as OTAS and NOTSS provide standardized expert-rating schemes for communication, leadership, coordination, and situational awareness.
- Socially grounded annotations remain scarce in operating rooms, motivating multilevel teamwork and counterfactual annotations for learnable agent feedback.
3 Source Dataset
The work builds on MM-OR, a multimodal operating-room dataset capturing realistic surgical interactions across synchronized visual, audio, robotic, system, and tracking streams. Its extension targets higher-level teamwork constructs absent from the source annotations.
- MM-OR contains 17 full-length surgical recordings of approximately 90 minutes each and 22 short clips ranging from 1 to 180 minutes.
- The dataset captures multiple sessions, days, and team compositions, providing variability in personnel, interactions, and workflow dynamics.
- Its modalities include ceiling-mounted RGB-D and RGB cameras, a low-exposure camera, wireless microphones, speech transcripts, robotic logs, screen recordings, procedural states, and 3D tracking.
- All modalities are hardware-synchronized for fine-grained multimodal and temporal analysis of surgical workflows.
- MM-OR supports perception and procedural tasks such as segmentation, tracking, scene graphs, phase recognition, next-action anticipation, and sterility-breach detection.
- The proposed extension addresses the source dataset’s lack of higher-level social and cognitive constructs by making teamwork the central benchmark focus.
4 Extended MM-OR
The extended MM-OR resource enriches surgical recordings with speaker- and language-aware representations, hierarchical teamwork annotations, and counterfactual judgments. Its multi-level design captures complementary aspects of team functioning across team, interaction, and individual perspectives.
- Extended multimodal resource: The resource adds speaker diarization, bilingual transcripts, hierarchical teamwork annotations, and counterfactual annotations to MM-OR.The enrichment combines interaction-centric representations with team-, interaction-, and individual-level labels.
- Annotation framework: Multiple annotation levels are used because no single framework or global score captures the dynamic interplay of communication, leadership, coordination, awareness, and role-dependent behavior.The framework separates team-level, interpersonal, and individual dimensions that provide partially independent information.
- Annotation protocol: Each six-minute clip was independently reviewed by three annotators using video, transcripts, and speaker-aware representations before counterfactual judgments were collected.The ordering first establishes descriptive ratings and then asks annotators to identify what may have driven the observed quality judgment.
- Interaction-level annotations: HMT ratings were centered on intermediate scores, with classes 2 and 3 comprising 63.6%–70.3% of samples across interaction dimensions.Communication and Joint Information Processing had nearly identical distributions.
- Individual-level annotations: NOTSS ratings were also concentrated at intermediate performance levels, with classes 2 and 3 forming majorities across Situation Awareness, Decision Making, Communication and Teamwork.The reported class shares were 64.7%, 72.1%, 57.4%, and 66.7%, respectively.
- Counterfactual layer: Counterfactual annotations identify actions or interactional moments judged to have most negatively affected teamwork and include rationales for alternative behavior.Communication failures account for approximately 28.0% of events, while monitoring or situational awareness issues account for 22.7% and cooperation or backup breakdowns for 20.0%.
5 Benchmarks
The benchmarks evaluate enriched multimodal representations, multi-level teamwork prediction, and counterfactual event identification. Across these use cases, temporal and relational modeling generally improves performance, while the experiments provide initial baselines rather than state-of-the-art claims.
- Benchmark scope: The benchmark suite covers multimodal feature enrichment, prediction of team-, interaction-, and individual-level constructs, and identification of counterfactual events linked to teamwork degradation.These use cases assess conversational representations, multi-level teamwork labels, and the counterfactual annotation layer.
- Feature enrichment: Speaker-aware and transcript-based representations substantially improve teamwork prediction over raw audio features, supporting conversational structure as important for modeling operating-room collaboration.The enrichment experiments progressively add structural, visual, audio, and textual information under LOGO evaluation across held-out surgical teams.
- Multi-level prediction: Meaningful predictive performance is achieved across team, individual, and interaction annotation layers, indicating that the annotated constructs are learnable from enriched multimodal observations.Results use macro F1-score, ten random seeds, and Leave-One-Group-Out evaluation across unseen surgical teams.
- Multi-level prediction: Temporal models generally outperform static baselines, while relational and spatio-temporal models provide further gains across teamwork prediction tasks.The reported pattern highlights the value of modeling temporal evolution and interpersonal relations explicitly.
- Counterfactual prediction: Counterfactual events can be operationalized as learnable prediction targets, with TE-ReNN achieving the strongest results by jointly modeling behavioral evolution and interpersonal interactions.This proof-of-concept connects qualitative expert reflections about alternative behaviors to computational prediction tasks.
6 Discussion
The resource extends operating-room data into multimodal, interaction-centered representations for analyzing surgical teamwork. Its discussion highlights practical limitations while positioning the dataset for integrated and explainable modeling.
- The dataset combines speaker-aware preprocessing, bilingual transcripts, multi-level teamwork annotations, and counterfactual representations for interaction-centered analysis.
- Expert and manual enrichment and annotation ensure high quality but limit scalability to larger datasets.The authors suggest semi-automated annotation pipelines as a possible direction for addressing this constraint.
- Integrating multiple teamwork frameworks preserves theoretical richness but introduces partial redundancy and potential inconsistencies across constructs.The heterogeneity may also support cross-framework alignment and latent structure discovery.
- Counterfactual annotations provide explanatory signals about perceived contributors to teamwork degradation but do not establish experimentally validated causal mechanisms.The annotations remain subjective because they reflect annotator interpretation.
- Privacy and accessibility constraints remain important boundaries for broader dissemination and reproducibility despite the dataset’s grounding in realistic operating-room recordings.
- The benchmark supports a shift from descriptive workflow modeling toward integrated, multimodal, and explainable teamwork modeling.