Source-linked AI summary

AffectSim: A Controllable Interactive 3D Simulation Benchmark for Embodied Affective Perception

Ke Xing, Zhilong Wang, Zheng Lian, Sicheng Zhao, Haifeng Lu, Zhen Zhang, Zitong Yu, Xiaojiang Peng, Changxin Huang, Runhao Zeng, Xiping Hu

arXiv:2608.25664v1cs.HC

TL;DR

Existing affective benchmarks generally fix observations before inference, leaving embodied evidence acquisition difficult to study systematically. AffectSim introduces replayable, factorized interactive 3D episodes with controllable observation, and its matched evaluations show that active sensing improves emotion-perception performance across most tested configurations. The benchmark provides an initial platform for studying controllable embodied affective perception, while remaining limited by acted behavior, simulation-to-reality gaps, restricted labels, and non-reactive simulated humans.

  • Problem

    Most affective benchmarks determine observations before inference, limiting systematic study of how embodied sensing influences affective perception.

  • Method

    AffectSim factorizes emotion-labeled motion from scenes and observation conditions in replayable interactive 3D episodes, evaluated with P-Init, P-Ref, and A-Obs protocols.

  • Results

    A-Obs improves 21 of 24 configurations, with mean Macro-F1 reaching 11.70% for open-source and 24.26% for closed-source models.

  • Takeaways & Limitations

    AffectSim makes affective observation controllable and provides an initial platform for repeatable, counterfactual studies of embodied affective perception.

  • Takeaways & Limitations

    The benchmark primarily uses acted expressions, has a simulation-to-reality gap, supports five discrete emotion categories, and uses non-reactive predefined human motions.

Abstract

from arXiv · show

Existing affective benchmarks largely consist of fixed recordings whose observation conditions are determined before inference, making it difficult to systematically study how embodied sensing influences affective perception. We introduce AffectSim, a controllable interactive 3D simulation benchmark for embodied affective perception. Rather than treating affective samples as fixed recordings, AffectSim instantiates emotion-expressive human motions as replayable 3D episodes in which distance, orientation, occlusion, scene geometry, and agent viewpoint can be systematically varied while preserving the underlying behavior and emotion label. AffectSim contains 27{,}647 episodes across five emotion categories and 57 scenes. Its factorized design separates affective behavior from observation conditions, supporting controlled re-observation of the same behavior as well as agent-controlled sensing in an executable 3D environment. To demonstrate this capability, we instantiate embodied emotion perception under matched initial (P-Init), reference (P-Ref), and actively acquired (A-Obs) observations. Across 24 frozen perception-model configurations, P-Ref substantially outperforms P-Init, while a simple two-stage active-observation baseline improves 21 of 24 configurations. Mean Macro-F1 increases from 9.89% to 11.70% for open-source models and from 22.61% to 24.26% for closed-source models, recovering 32.0% and 20.1% of their respective P-Ref--P-Init gaps. Episode-level recovery and path-aware evaluation further characterize the current baseline beyond aggregate recognition performance. These results demonstrate the value of making affective observation controllable and establish AffectSim as an initial platform for studying embodied affective perception through interactive 3D simulation.

1 Introduction

Existing affective benchmarks largely fix observations before inference, limiting controlled study of how embodied sensing affects emotion perception. AffectSim addresses this gap with replayable, factorized 3D episodes and matched protocols for measuring and recovering observation-quality differences.

  • Most affective benchmarks evaluate prerecorded observations whose distance, orientation, and visibility are determined before inference.
  • AffectSim turns observation into a controllable experimental variable by replaying the same affective behavior while varying scene and sensing conditions.Its factorized design separates affective motion from scenes and observation conditions.
  • 27,647 episodes span 57 scenes, with controlled distance, orientation, occlusion, navigation, and agent-initialization conditions.The construction includes single-person and dyadic episodes generated from reusable motion assets.
  • P-Init, P-Ref, and A-Obs measure initial evidence, privileged reference evidence, and actively acquired evidence from the same initial state.The observation policy is decoupled from the recognizer, enabling evaluation across perception models.
  • 21 of 24 model configurations improve with active observation, reaching mean Macro-F1 scores of 11.70% for open-source and 24.26% for closed-source models.These recover 32.0% and 20.1% of the respective P-Ref–P-Init gaps.
  • AffectSim focuses on controllable perceptual observation, while affective behavior is prescribed rather than generated interactively.The benchmark is positioned as a complement to real-world affective datasets and a platform for repeatable, counterfactual experiments.

2 Related Work

Related affective benchmarks increasingly provide richer evidence and embodied settings, but usually present that evidence before inference. AffectSim instead places re-observation and agent-controlled sensing inside an interactive 3D evaluation loop.

  • Emotion understanding from fixed observations: Conventional affective benchmarks provide increasingly rich facial, bodily, audiovisual, contextual, and reasoning evidence, but observations remain predetermined.
  • AffectSim’s distinction: Its construction pipeline converts emotion-labeled performances into reusable motion assets and composes them with scenes, placements, initializations, and observation challenges.
  • Embodied and egocentric affective computing: Robot-centric and egocentric benchmarks add embodied contexts, yet typically evaluate interpretation of supplied evidence rather than its physical acquisition.
  • AffectSim’s distinction: AffectSim combines interactive 3D simulation, controlled re-observation of the same affective event, and agent-controlled observation in one framework.Agent actions can change the visual evidence subsequently available for emotion recognition.

3 AffectSim Design and Construction

AffectSim constructs replayable, factorized 3D episodes that separate affective behavior from controllable observation conditions. Its modular pipeline supports diverse motion assets, scenes, embodied challenges, validation, and repeated observation of fixed behavior.

  • Benchmark construction: AffectSim converts emotion-labeled performances into reusable motion assets, composes them with scenes and embodied observation settings, then validates generated episodes.The three-stage construction pipeline supports adding motions, scenes, and observation settings without redefining the task.
  • Behavior–observation decoupling: AffectSim represents episodes as configurations combining motion, scene, human placement, agent initialization, and observation challenge, while inheriting emotion labels from the motion.This factorization allows distance, orientation, occlusion, scene layout, obstacles, and initialization to vary while behavior and labels remain fixed.
  • Replayable episodes: Replayable episodes provide egocentric RGB-D and proprioceptive observations, allowing agents to alter future observations through movement and supporting matched reference, initial, monitor, and active observations.The asset–episode separation also supports extensibility to new motions, performers, dyadic behaviors, avatars, scenes, and observation challenges.
  • Affective motion assets: The shared label space contains angry, fearful, happy, neutral, and sad, excluding categories without clear cross-dataset correspondence rather than force-merging them.Motion assets combine KDAE, Emilya, and a newly captured 48-sequence dyadic pilot dataset.
  • Embodied episode generation: Observation-controlled composition varies far-visible, rear-view, partial-occlusion, and obstacle-detour challenges without changing the expressed emotion.The dyadic pilot stores both participants while designating one as the recognition target and yields 1,440 composed episodes.
  • Benchmark composition: 27,647 episodes span five approximately balanced emotion categories and diverse scenes, affective motions, action contexts, distances, orientations, occlusions, and navigation difficulties.The benchmark includes both single-person and dyadic interactions, with representative motion assets covering the five emotion categories.

4 Evaluation Protocol

AffectSim evaluates embodied emotion recognition under matched passive and active observation protocols, separating recognition quality from how visual evidence is acquired. The protocol measures observation benefit, episode-level recovery, path efficiency, and acquisition costs across group-disjoint episodes.

  • Observation Protocols: P-Init uses the agent’s stationary initial camera pose, while P-Ref supplies a reproducible high-quality trajectory generated with privileged geometry.P-Ref is a diagnostic reference rather than an oracle or upper bound.
  • Observation Protocols: A-Obs starts from the P-Init state and lets an observation policy move, accumulate views, and terminate acquisition before recognition.The same policy is applied across recognizers, and invalid acquisition falls back to the paired P-Init observation.
  • Evaluation Dimensions: Macro-F1 is the primary recognition metric because the five emotion categories are not perfectly balanced, with per-class recall as an additional diagnostic.Passive settings use predefined observations, whereas active evaluation tests whether movement can acquire different visual evidence.
  • Evaluation Dimensions: OGR quantifies the fraction of the P-Ref–P-Init recognition gap recovered through active observation, allowing values above one and negative values.It is reported only when P-Ref exceeds P-Init, with OGR = 0 denoting no improvement and OGR = 1 matching P-Ref.
  • Evaluation Dimensions: Recovery Rate measures how often active observation corrects an initially wrong prediction at the episode level, complementing aggregate metric gains.E-SPL additionally combines final recognition success with path efficiency at reference distances of 1, 3, and 5 m.
  • Evaluation Dimensions: AffectSim separately reports total path length, budget exhaustion, and clipped moves to characterize where active policies spend movement budgets and fail acquisition.Splits keep variants from the same source performance together; shared scenes across splits therefore test unseen performers and source motions, not unseen environments.

5 Experiments

Experiments evaluate embodied emotion recognition under matched observation protocols using frozen models, showing that observation quality strongly affects recognition and that active sensing recovers part of the gap.

  • Experimental Setup: The evaluation uses paired episodes with unchanged scene, motion, and emotion label while varying observation protocols and frozen recognizers.P-Init uses the initial pose, P-Ref uses a privileged reference trajectory, and A-Obs actively changes viewpoint before recognition.
  • Experimental Setup: A two-stage active-observation baseline combines environment search and handoff with target following based on person location and depth cues.The policy does not use emotion labels, oracle trajectories, simulator masks, or recognizer outputs.
  • Effect of Observation Quality: P-Ref raises mean Macro-F1 from 9.89% to 15.56% for open-source models and from 22.61% to 30.78% for closed-source models.P-Ref achieves the highest score for 17 of 19 open-source configurations and 22 of 24 configurations overall.
  • Effect of Observation Quality: P-Mon reaches 10.31% Macro-F1 for open-source models, below P-Ref’s 15.56%, showing that an external room-level view alone does not reproduce informative observation.Observation quality depends jointly on visibility, framing, orientation, distance, and exposed motion cues.
  • Recognition Gains from Active Observation: A-Obs outperforms P-Init for 21 of 24 configurations, increasing mean Macro-F1 from 9.89% to 11.70% for open-source and from 22.61% to 24.26% for closed-source models.On the new-observation subset, gains are 3.40 and 2.91 percentage points, compared with 1.81 and 1.64 points on the complete protocol.
  • Recovery and Acquisition Efficiency: A-Obs recovers 32.0% of the P-Ref–P-Init gap for open-source models and 20.1% for closed-source models, leaving a substantial gap to P-Ref.Episode-level Recovery Rate averages 3.10% for open-source and 7.24% for closed-source models; mean E-SPL is 1.70%, 0.95%, and 0.78% at 1, 3, and 5 m.
  • Diagnostics: Active-observation benefits vary across emotion classes and episode families rather than being uniform.For open-source models, sad and happy show the largest recall gains while fearful receives little improvement; episode families also exhibit different performance patterns.
  • Recovery and Acquisition Efficiency: 47.73% of episodes reach stable handoff, while 52.27% terminate after exhausting the execution budget.The baseline averages 23.23 m per episode, with 97.4% of locomotion before handoff and 14.10% of attempted translations clipped by the NavMesh.

6 Conclusion

AffectSim makes affective observation controllable by replaying emotion-labeled human motion in executable 3D environments. Experiments show active observation improves many frozen recognizers but remains incomplete, while the benchmark is limited by acted behavior, simulation-to-reality gaps, discrete labels, and nonreactive virtual humans.

  • 6 Conclusion: AffectSim instantiates emotion-labeled human motions as replayable events while systematically controlling scene and observation conditions through agent actions.The design separates affective motion from scene and observation conditions, preserving the underlying behavior and emotion label during re-observation.
  • 6 Conclusion: Across 24 frozen configurations, P-Init, P-Ref, and A-Obs reveal a clear observation-quality difference, while A-Obs improves 21 of 24 configurations.Episode-level recovery and path-aware evaluation complement aggregate recognition performance.
  • 6 Conclusion: AffectSim provides an initial platform for repeatable and counterfactual studies of affective perception in controllable interactive 3D environments.The current benchmark focuses on prescribed, replayable affective behavior and controllable observation rather than fully closed-loop interaction.
  • 6.1 Limitations: The benchmark primarily uses acted expressions, has a simulation-to-reality gap, covers five discrete emotion categories, and uses predefined nonreactive human motions.These boundaries focus the benchmark on embodied affective perception under controllable observation rather than fully closed-loop affective interaction.
  • 6.2 Ethics and Broader Impact: Emotion inference may be misused for surveillance, manipulation, or psychological profiling.The benchmark is intended for research rather than high-stakes decision making or surveillance.
Loading 2608.25664v1…