Source-linked AI summary

Scales, Reflections, and Conversations: A Multi-Modal Approach to Emotion Annotation

Pragya Singh, Prashasti Gupta, Hitesh Bhandari, Kanishk Goel, Mohan Kumar, Pushpendra Singh

arXiv:2609.05046v1cs.HC

TL;DR

Existing emotion-data collection often relies on fixed prompts and single-format reports that can miss participants’ availability, agency, and expressive complexity. The paper evaluates a participant-centric EMA application combining flexible scheduling with multimodal logging, finding that modality choice supports richer and more situated emotion data while sustained engagement remains challenging.

  • Problem

    Existing emotion-reporting systems can provide limited contextual and expressive information because they often rely on fixed schedules and single modalities.

  • Method

    The paper conducts a formative one-week field study of a participant-centric EMA prototype with flexible prompting schedules and multiple emotion-logging modalities.

  • Results

    Multimodal logging enabled users to adapt expression to situational constraints and cognitive load, supporting richer and more nuanced emotion data.

  • Takeaways & Limitations

    Choice-based, user-adaptive emotion logging is positioned as a promising approach for capturing heterogeneous, situated everyday emotional experiences.

  • Takeaways & Limitations

    The formative study involved 33 predominantly well-educated, technologically proficient participants from a similar cultural context over one week, limiting generalizability.

Abstract

from arXiv · show

Mental health concerns are increasing worldwide, highlighting the need for interventions that support everyday emotional well being. Prior work has demonstrated the potential of wearable and mobile technologies to deliver data driven interventions. However, developing effective data-driven systems requires access to emotion data that captures individuals' emotional variability and change in everyday contexts. Existing approaches to data collection largely rely on frequent, prescheduled prompts and predefined scales or questionnaires. These methods often fail to account for participants' availability, agency, or the complexity of their emotional experiences, resulting in shallow, context poor data. In this paper, we present a feasibility study of a participant centric, multimodal emotion-annotation application designed around users' emotional intensity and availability. Our findings show how multimodal emotion logging can shape participants' experiences and data logging behaviors, and demonstrate its potential to support the collection of richer, more nuanced emotion data.

1 Introduction

Emotion self-reports are essential for interpreting indirect behavioral and physiological proxies, but existing systems struggle to collect nuanced, contextual data without overburdening participants. This paper studies a participant-centric EMA system that combines flexible prompting with multimodal emotion logging.

  • Self-reports are needed because observable behavioral signals have an indirect relationship with subjective emotional experiences.
  • Frequent reporting can disrupt daily routines, while limited contextual awareness and expression options reduce the richness and relevance of emotion data.
  • Single-modality EMA approaches often assume one interaction format can capture emotional experience despite variation across situations and individuals.
  • The study presents a multimodal, participant-centric prototype designed to balance user burden, expressive flexibility, and scalable real-world deployment.
  • A seven-day feasibility study with 33 participants examines engagement, expressive elaboration, and contextual grounding in multimodal emotion logging.
  • Findings suggest that modality switching helps users adapt emotional expression to situational constraints and cognitive load.

2 Related Work

Emotion self-reporting has progressed from standardized scales toward more naturalistic and interactive approaches, but existing methods still face contextual and expressive limitations. The reviewed work motivates systems that better support participant-centered, in-the-wild reporting.

  • Categorical and dimensional theories have shaped predefined emotion labels and affective scales used in self-reporting.
  • Many EMA tools provide limited contextual information, which makes emotional changes difficult to interpret for predictive modeling.
  • Interactive systems have introduced reflective writing, opportunistic reporting, gamification, AI-generated content, and passive-data-supported reflection.
  • Naturalistic and semi-naturalistic datasets combine physiological, behavioral, and self-report data across everyday contexts.

3 Application Design: Overview

The application combines user-controlled scheduling, on-demand logging, and three emotion-reporting modalities to accommodate different contexts, cognitive loads, and expressive needs. Its structured and reflective pathways add contextual detail while preserving flexible participation.

  • Users can define reminder times or log emotions on demand, reducing reliance on externally imposed schedules.
  • Across the modes, users can select the reporting format that fits their context, cognitive load, and preferred level of expression.
  • The home screen presents four scheduled reporting slots and an on-demand action button for emotion logging.
  • Quick Mode: Quick Mode combines arousal–valence quadrant selection, conditional stress measurement, multiple emotion labels, contextual activities, and confidence ratings.
  • Detailed Reflections: Detailed Reflections use optional prompts and multimedia inputs to support complex, context-rich emotional experiences without requiring open-ended reflection.
  • LLM-supported Annotations: An LLM-supported conversational interface provides interactive scaffolding for emotions that are ambiguous, evolving, or difficult to articulate.

4 Feasibility Study

The feasibility study contextualized participants’ emotion-logging behavior through surveys, onboarding, field use, and post-study feedback. It examined engagement across participant backgrounds and everyday reporting contexts.

  • The pre-study survey collected demographics, mental health history, routines, emotional events, and comfort with emotional expression.
  • Participants used the application for one week after onboarding that covered installation, reporting procedures, and privacy practices.
  • The study analyzed collected data alongside feedback surveys and semi-structured exit interviews to assess engagement, usability, relevance, and satisfaction.
  • Recruitment used snowball and convenience sampling; 35 enrolled and 33 completed the study without paid incentives.
  • Participant characteristics included demographic, psychological, clinical, contextual, and emotional-attitude measures relevant to interpreting engagement.

5 Analysis

The study used mixed-methods analysis to examine how scheduling flexibility, multimodality, emotion selection, and media sharing shape everyday emotion logging and expressive characteristics. Exploratory mixed-effects models accounted for repeated observations and participant-level variability.

  • The study combined descriptive quantitative analysis, mixed-effects modeling, and qualitative interpretation to examine behavior and expressive characteristics in everyday emotion logging.
  • The exploratory analysis examined how impromptu versus scheduled logging, modality choice, and individual characteristics relate to flexibility, expressive elaboration, and interaction patterns.
  • Mixed-effects models treated participant as a random effect to account for repeated measures and inter-individual variability in logging behavior.
  • The models specified scheduling type and annotation modality as fixed effects, with dependent variables tailored to each exploratory question.

6 Findings

Findings indicate that participants engaged more with impromptu than scheduled emotion logging, while combining both approaches supported different reporting needs. Participants generally favored quick, flexible annotations, but richer modalities remained useful when emotions were stronger or harder to interpret.

  • Engagement by scheduling: Scheduled prompts were associated with significantly lower response probabilities than impromptu opportunities (β = −0.724, SE = 0.027, z = −26.99, p < .001).The pattern was consistent with observed response rates and reflected a substantial practical difference between scheduling conditions.
  • Engagement by scheduling: Participants valued combining prescheduled notifications with self-initiated logging, using flexibility to record emotions when they recognized a need.Participants also suggested changing prescheduled slots weekly or daily to accommodate varying routines.
  • Engagement by scheduling: Evening slots were most popular (35.61%), followed by afternoon slots (31.82%), while early-morning slots were least selected (4.55%).Participants favored notifications during active hours from midday through evening, with minimal interest in early-morning and late-night interruptions.
  • Modality use: Quick mode accounted for 433 of 505 emotion logs (85.7%), compared with 52 detailed entries (10.3%) and 20 LLM-assisted logs (4.0%).LLM annotation data for four participants were not correctly recorded, potentially underestimating engagement in that condition.
  • Modality use: Modality choice had minimal influence on the structural characteristics of emotional reporting, with participant-level differences contributing more variance overall.The modality effect was small (marginal R2 = .004), whereas participant-level differences accounted for a larger proportion of total variance (conditional R2 = .487).
  • Modality use: Participants reported using quick annotations for routine or less intense experiences and detailed reflection when they had more time or struggled to identify their emotions.Self-initiated logs were also preferred for negatively charged emotions, while prescheduled entries showed a higher positive-to-negative ratio (2.71 vs. 1.80).

7 Discussion

The feasibility study found that multimodal logging primarily changed the expressive depth and interpretability of reports rather than the range of emotions reported. Richer modalities supported contextualization and reflection but introduced time-cost and interaction-quality tensions.

  • 7 Discussion: Multimodal logging primarily changed how emotions were expressed and interpreted, not the underlying range of emotions reported.Quick-mode entries captured a broad distribution of emotional states, while journal and conversational entries added elaboration and context.
  • 7 Discussion: Structured quick reporting remained sufficient for capturing everyday emotional states in ecological settings.Multi-select labels captured co-occurring emotions, and activity lists provided basic contextualization.
  • 7 Discussion: Journal and conversational entries made categorical labels more interpretable by adding layered accounts and contextual narration.The same label could reflect different physical, mental, or situational experiences.
  • 7 Discussion: Conversational logging often functioned as a reflective scaffold, helping participants unpack and make sense of vague or compressed emotional labels.Participants used conversation for reflection, although some found it time-consuming, repetitive, or insufficiently human-like.
  • 7 Discussion: Conversational richness involved a design trade-off between deeper emotional elaboration and participants’ limited time or availability.Participants reported higher time costs when emotions were straightforward or they had limited availability.

8 Limitations

The study’s one-week, 33-participant design and relatively narrow sample constrain generalization to broader populations and longer-term deployments. Engagement may also have been shaped by privacy concerns, lack of incentives, and seasonal context.

  • 8 Limitations: The one-week study with 33 participants limits generalization to broader populations and longer-term deployment contexts.Participants were primarily well-educated, technologically proficient, and drawn from a similar cultural context.
  • 8 Limitations: The sample may not represent people with lower digital literacy, different backgrounds, diverse cultural contexts, or more severe mental health conditions.The authors recommend evaluation with more diverse populations over longer periods.
  • 8 Limitations: Some participants hesitated to share emotions despite privacy measures, potentially reducing data richness.The study used private-server deployment, anonymization, and secure storage.
  • 8 Limitations: Unincentivized participation may have favored people already comfortable with emotional self-reflection or intrinsically motivated to engage.Observed engagement may therefore not generalize to people requiring stronger external motivation.
  • 8 Limitations: Overlap with the festival season influenced user engagement and highlighted the importance of situational context.

9 Conclusion

The paper introduces a choice-based multimodal annotation system for transient emotions, allowing people to log experiences according to emotional intensity and availability. The feasibility findings suggest that this approach can support richer, more nuanced emotional data while emphasizing warm, user-led interaction.

  • 9 Conclusion: The system uses a choice-based design so users can log transient emotions according to their current intensity and availability.
  • 9 Conclusion: Multimodal emotion logging can support richer and more nuanced emotional data across diverse and dynamic emotional profiles.
  • 9 Conclusion: The journaling assistant is designed for Indian users aged 18–60 and encourages natural, comfortable, judgment-free self-expression.
  • 9 Conclusion: The interaction approach accommodates expressive and reserved users by mirroring tone, using warmth, and avoiding imposed emotional interpretations.
  • 9 Conclusion: Users are meant to control conversational depth, with gentle invitations to elaborate rather than pressure toward prolonged introspection.

B Pre-Study Survey

The pre-study survey examined participants’ mental-health history, daily-life context, emotional-expression comfort, and concerns about others’ perceptions. These measures characterized factors relevant to emotion-reporting behavior before the study.

  • B Pre-Study Survey: The survey asked about diagnosed mental-health conditions, counseling or therapy, recent emotional events, and prior use of emotion-tracking applications.
  • B Pre-Study Survey: Participants reported how structured their daily routines were and how supportive or difficult their family dynamics felt.
  • B Pre-Study Survey: Participants rated their ability to balance responsibilities and personal time across a five-level response scale.
  • B Pre-Study Survey: The survey assessed comfort expressing emotions, concern about others’ perceptions, and whether emotional expression was viewed as a sign of weakness.

C Interview Questions

The interview protocol examined participants’ experiences, logging-method preferences, emotional reflection, privacy concerns, satisfaction, and desired improvements.

  • Participants were asked to describe their overall experience, standout aspects, prior experience with similar applications, and changes over the study week.
  • The protocol compared three emotion-logging methods: the arousal–valence quadrant, chatbot interaction, and guided audio-and-image prompts.
  • Questions examined how the application affected daily routines, emotional processing, behavior, decision-making, and perceived emotional patterns.
  • The evaluation covered tutorial helpfulness, notification timing, recording-option use, ease of use, and effectiveness in capturing emotions.
  • Participants were asked about privacy, comfort sharing emotions, annotation challenges, and the depth and focus of their emotional assessments.
  • Satisfaction questions addressed daily-life usability, continued-use intentions, helpful features, desired improvements, and final suggestions.

E Technical Implementation

The application used a cross-platform mobile stack, Firebase-backed storage and authentication, notifications, and a locally deployed, privacy-oriented chatbot evaluated across emotional states.

  • The application was developed with React Native and Expo for deployment on Android and iOS.
  • Firebase provided Firestore for textual data, Firebase Storage for audio and images, and authentication for user management.
  • Notifications used the Notifee library in React Native to prompt users for emotion annotations.
  • The chatbot was deployed locally using LLaMA 3.3 70B Instruct, fine-tuned on counseling-oriented HOPE dialogues with LoRA and 4-bit quantization.
  • Chatbot guidance emphasized empathetic journaling, cultural sensitivity, user-led reflection, validation, and gentle conversational closure.
  • Responsiveness across emotional states was assessed using eight Russell Circumplex Model-aligned texts and three user types.

F Additional Information

Additional materials document participants’ scheduled and impromptu logging, tutorial screens, annotation-confidence responses, and the application’s emotion list.

  • 146 scheduled prompts and 359 impromptu logs were completed by 33 participants across the study period.The reported means were 4.4 scheduled prompts and 10.9 impromptu logs per participant.
  • The prompt-completion figure presents individual counts of scheduled and impromptu prompts for each participant.
  • Tutorial screens introduced annotations and explained the valence–arousal quadrant using examples.
  • Table 7 reports annotation confidence and activity responses across all user entries (N=505).
  • Table 8 lists the emotions used in the application.
Loading 2609.05046v1…