Source-linked AI summary

Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration

Youngseok Seo, Sueun Jang, Hyesoo Park, Renz Samuel Gutierrez, Joseph Seering, Uichin Lee

arXiv:2609.11529v1cs.HC

TL;DR

Ethics education for STEM students needs scalable ways to support collaborative reflection on societal consequences, beyond largely manual and individual-focused approaches. Ethics Training Agents provides a structured human-AI group discussion using a facilitator and role-based LLM participants, and a study with 45 undergraduates found enhanced ethical sensitivity, engagement, coordination, and perspective-taking alongside important social and methodological limitations.

  • Problem

    STEM ethics education faces scalability challenges, while existing approaches often provide limited collaborative deliberation and inconsistent group learning experiences.

  • Method

    The system combines Judgment Call-based role play with an LLM facilitator and multiple LLM agents embodying distinct ethical perspectives.

  • Results

    The 45-participant study found enhanced ethical sensitivity, sustained engagement, coordination, and perspective-taking, while agents also lowered social barriers at the cost of shallower human-human interaction.

  • Takeaways & Limitations

    Consistent facilitation and topic alignment can support coordination and reflection, while richer debate may require less agreeable agents and more diverse argumentative strategies.

  • Takeaways & Limitations

    Because the evaluation used only pre–post comparisons without a control condition, the observed effects cannot be causally separated from novelty or general learning effects.

Abstract

from arXiv · show

Group-based ethics training for Science, Technology, Engineering and Mathematics (STEM) students is a complex challenge, requiring substantial resources and expertise. While activity-based teaching methods, such as role-playing and discussions, are commonly employed to simulate real-world scenarios, current practices are often manual and lack integration with effective online platforms for supporting group-based ethical discussions. In this work, we propose Ethics Training Agents, a group discussion system that leverages multiple LLM participants embodying distinct ethical orientations, along with a moderator agent, to enable structured human-AI group ethical discussions for collaborative reflection. We conduct a user study with 45 undergraduate STEM students to evaluate the learning outcomes and user experience. The results show that our system supports engagement, coordination, and perspective-taking in group discussions and has a positive influence on ethical sensitivity. We also discuss practical design strategies for integrating multiple LLM agents into multi-human group settings to facilitate ethics training for STEM students.

1 Introduction

STEM students need early ethics education because engineering decisions affect society, yet group-based role-play activities are difficult to scale. Ethics Training Agents addresses this gap with multiple role-based LLM participants and a facilitator for collaborative ethical discussion.

  • Engineering technologies affect safety, privacy, equity, and sustainability, while STEM curricula often leave students underprepared for their societal consequences.
  • Role-playing and discussion can support stakeholder perspective-taking and surface ethical concerns in realistic group activities.
  • Scaling role-play ethics education is difficult because classroom groups require substantial instructor involvement and facilitation expertise.
  • Existing digital ethics tools support reflection and multiple perspectives but remain largely unidirectional and individual-focused, leaving limited support for collaborative deliberation.
  • Ethics Training Agents uses an LLM facilitator and multiple LLM participants with distinct ethical backgrounds to reduce facilitation overhead and provide consistent perspectives across groups.
  • The study examines ethical sensitivity, facilitation of role-play discussions, perceived human-versus-agent contributions, and social dynamics.

2 Related Work

Prior ethics education and digital tools offer frameworks, role-play, and guided reflection, but often provide limited support for scalable, multi-party deliberation. This work adapts role-playing and discussion into a multi-agent system for more consistent collaborative ethics education.

  • Ethics education draws on philosophical frameworks, value sensitive design, and principle-based exercises such as fairness, transparency, accountability, inclusivity, safety, and privacy.
  • Existing ethics education includes dedicated courses, embedded modules, discussions, debates, and oral arguments, but modular approaches face scalability challenges.
  • Role play can immerse students in realistic scenarios and stakeholder perspectives, yet requires substantial facilitation, time, coordination, and consistency across groups.
  • The paper addresses these gaps through a module-level, discussion-based approach using multiple LLM agents with distinct ethical and philosophical backgrounds.
  • Fictional scenarios, role play, and game-based methods are established approaches for engaging learners in ethical reflection and exploring future harms.
  • Digital tools such as PEaRCE, AI LEGO, and Responsible and Inclusive Cards support individual reflection or artifact-based analysis but offer limited interactive, multi-party deliberation.

3 System Design

The system adapts Judgment Call into a roughly 75-minute, staged group discussion with structured phases, an LLM facilitator, and persona-based agents. The design also includes iterative testing and safety oversight.

  • Judgment Call as a Foundation Structure.: The system uses Judgment Call as a team-based, discussion-oriented foundation designed around VSD, explicit objectives, and a single class session of approximately 75 minutes.
  • Judgment Call as a Foundation Structure.: Participants identify stakeholders, write role-based reviews, identify ethical problems, and propose technical, policy, or social solutions.
  • Discussion Structure: Each discussion stage uses idea-gathering, deliberation, questioning, and wrap-up to move from diverse ideas toward selected shared inputs for later stages.
  • Facilitator Design: The LLM facilitator controls flow, manages speaking turns and time, summarizes discussion content, and selects ideas for subsequent stages.
  • Iterative Design: Iterative pilot testing refined the interface, prompts, and discussion structure, while trained facilitators remained available because appropriate agent behavior could not be guaranteed.

4 Implementation

The implementation separates the user interface, synchronized discussion state, and agent behavior orchestration into three modules represented as explicit discussion states.

  • The system comprises a web client, a WebSocket server, and an agent orchestration module that manages behavior according to the discussion structure.
  • Each discussion stage and phase is implemented as an explicit state node using the LangGraph framework.

5 User Study Methods

The study evaluated a classroom-based group discussion system with undergraduate STEM students using a Black Mirror design-fiction scenario, pre/post ethical-sensitivity measures, peer ratings, discussion logs, and interviews. Quantitative analyses included reliability checks and statistical tests suited to the survey and peer-rating data.

  • Participants and Procedure: 45 undergraduate STEM students participated in 15 groups of three during an approximately two-hour classroom study.Freshmen were excluded; participants averaged 22.1 years of age, with 42% female and 58% male.
  • Discussion Scenario: Participants role-played engineers designing a smart lens with a miniature camera for memory augmentation, based on a Black Mirror scenario.The shared scenario provided a consistent foundation for randomly matched groups.
  • Measures and Data Collection: The evaluation combined pre/post ethical-sensitivity scores, post-discussion peer ratings, interaction logs, and focus-group interviews.These measures addressed educational impact, perceptions of human and LLM agents, social interaction, and participant reflections.
  • Measures and Data Collection: Ethical sensitivity was operationalized as ethical issue awareness and ethical consequence reasoning in technology-design situations.The constructs capture recognizing ethically salient features and reasoning about consequences and alternative actions.
  • Measures and Data Collection: Peer ratings assessed contribution, diversity, and influence to compare participants’ perceived roles in the discussion.Contribution ratings were adapted from a peer-evaluation framework for group work, while additional ratings assessed persona embodiment.
  • Data Analysis: The analysis used reflexive thematic analysis for interviews, reliability checks for quantitative scales, t-tests and Wilcoxon tests for survey measures, and Kruskal–Wallis tests for peer ratings.Peer-rating measures failed normality tests, motivating the nonparametric group comparison; all reported scales had acceptable internal consistency.

6 Results

The system improved ethical sensitivity while supporting structured, time-bounded discussion, perspective-taking, and reflection. Participants valued facilitation and diverse viewpoints, but reported trade-offs involving discussion depth, agent acceptance, and human–human interaction.

  • 6.1 RQ1. Changes in Ethical Sensitivity in Technology Design: Pre-post tests showed significant gains in ethical issue awareness and ethical consequence reasoning, with awareness improving more substantially.Participants attributed awareness gains especially to Identify Stakeholders and Write Reviews, which broadened their consideration of people and situations.
  • 6.2 RQ2. Facilitating Structured Ethical Discussions: Participants found the divergence, groan, and convergence structure easy to follow and suitable for covering diverse ethical perspectives within limited time.The structure supported brainstorming and parallel exploration, although turn-taking could limit in-depth discussion of individual ideas.
  • 6.2 RQ2. Facilitating Structured Ethical Discussions: The LLM facilitator supported smooth completion by coordinating speaking turns, managing time, maintaining coherence, and summarizing discussion content.Participants reported that summaries helped them catch up when the conversation moved too quickly or became difficult to follow.
  • 6.2 RQ2. Facilitating Structured Ethical Discussions: LLM participants clarified stage goals, kept discussion on topic, and prompted users to articulate and examine ethical concerns more critically.Their questions helped participants move from broad labels toward specific concerns, such as potential misuse of user information.
  • 6.2.4 LLM Participants Facilitating Diverse Perspective Taking: Distinct LLM viewpoints expanded perspective-taking beyond STEM-only groups and consistently surfaced minority or less dominant concerns.Participants said these perspectives appeared without the social pressures that can discourage human minority viewpoints.
  • 6.3 Discussion of LLM Contributions: Human peers were rated significantly more positively than every LLM persona on general contribution, diversity, and influence, while LLM personas did not differ significantly from one another.Participants also reported that breadth of AI-generated ideas could come at the expense of depth, and that agents’ sycophantic responses lowered perceived contribution.

7 Discussion

The discussion identifies how multi-agent ethics education can support perspective-taking and structured participation while exposing trade-offs in reasoning depth, human interaction, trust, and representation. It proposes process-oriented, transparent, and more challenging agents, alongside safeguards and study designs that address these limitations.

  • 7.1 Beyond Ideas Toward Ways of Thinking: Learning Ethical Reasoning Through Multi-Agent AI Interactions: Participants wanted agents to explain how ethical judgments are formed, not merely generate additional ideas.They sought ethical mentors and peers, but persona prompts focused mainly on stage-specific ideas and did not sufficiently expose reasoning frameworks.
  • 7.1 Beyond Ideas Toward Ways of Thinking: Learning Ethical Reasoning Through Multi-Agent AI Interactions: Process-oriented agent reasoning should be paired with safeguards because fluent explanations may encourage over-reliance and erode learner agency.The system positioned agents as undergraduate peers and included a Questioning phase so participants could challenge them.
  • 7.2 Managing Expectation and Identity Disclosure to Mitigate AI Agent Dismissal: Participants often treated agent errors as generative-AI defects rather than bounded persona limitations, quickly discrediting the persona.Expectation-setting and disclosure of role-specific limitations were proposed to frame shortcomings as part of the learning experience.
  • 7.3 Designing for Balanced Social Dynamics in Mixed Human–AI Ethical Discussions: Agents’ diverse perspectives broadened exploration but sometimes reduced discussion depth and critical challenge.Participants described divergent idea generation, while overly agreeable responses and limited lived experience made discussions shallower.
  • 7.3 Designing for Balanced Social Dynamics in Mixed Human–AI Ethical Discussions: Human–AI discussion outcomes depend on response strategy, persona design, and mechanisms that preserve human–human interaction.Suggested interventions include counterarguments, richer simulated experiences, peer-response requirements, and explicit reminders that agents lack lived experience.
  • 7.4 Limitations and Future Work: Designed personas cannot represent all relevant ethical positions, and fluent unsupported responses may still produce bias or over-reliance.The authors recommend treating agents as contestable prompts rather than sources of a correct ethical answer.
  • 7.4 Limitations and Future Work: The study’s conclusions are constrained by pre–post evaluation without a control condition, configuration choices, short-term measurement, and unfamiliar group members.These factors limit causal interpretation, long-term conclusions, and generalization to established classroom relationships.

8 Conclusion

The paper presents Ethics Training Agents as a multi-agent system for collaborative STEM ethics training. In a study with 45 participants, it supported ethical sensitivity, engagement, coordination, and perspective-taking while revealing design needs for richer debate and careful management of agent expectations.

  • 8 Conclusion: The study with 45 participants found enhanced ethical sensitivity, sustained engagement, and perspective-taking in mixed human–AI group discussions.It also identified social dynamics and design implications for future multi-agent discussion systems.
  • 8 Conclusion: Consistent facilitation and topic alignment promoted coordination and reflection, while richer debate may require less agreeable agents and more diverse argumentative strategies.The conclusion frames these as design implications for future education systems.
  • 8 Conclusion: Participants held ambivalent expectations toward agents combining extensive knowledge with distinct persona identities.The conclusion links these expectations to the goal of fostering deeper ethical reasoning.

A Appendix: Measurement Instruments and Peer Rating Analysis

The appendix documents the study’s measurement instruments and the statistical analysis used for peer ratings. It includes the ethical-sensitivity scale, peer-rating dimensions, and Kruskal–Wallis results with post-hoc comparisons.

  • The appendix provides the Ethical Sensitivity in Technology Design instrument used in the study.
  • The peer-rating instrument defines dimensions and items for evaluating participants in group discussions.
  • Table 3 reports non-parametric Kruskal–Wallis results as mean (SD), with Holm-corrected Dunn comparisons against human peers.
  • No significant differences were observed among the AI agents, while comparisons against human peers are reported separately.

B Appendix: Program Structure

The appendix describes the system implementation and presents its overall architecture alongside the orchestration module’s node structure.

  • Figure 5 presents the overall system architecture and the node structure of the agent orchestration module.
  • The appendix documents the system implementation and provides access to the authors’ code repository.
  • The implementation description links the architecture to the orchestration logic used to run the discussion system.

B.1 Architecture Overview

The system uses separate client, server, and agent-orchestration modules. The client handles interaction, the server synchronizes shared state, and the orchestration module controls discussion procedure.

  • The system is organized into three modules: client, server, and agent orchestration.
  • The client renders the discussion interface and transmits participant actions to the server.
  • The server relays events and ensures that human participants and LLM agents share the same state.
  • Built on LangGraph, the agent orchestration module makes decisions governing the discussion procedure.

B.2 Server and Communication Layer

The communication layer uses WebSocket-based real-time updates so participants can observe actions and maintain a synchronized discussion view.

  • WebSocket communication supports real-time observation of hand raising, speaking-turn transfers, and incoming utterances.
  • The server broadcasts participant or agent events, including hand raises, utterances, and speaking-turn passes.
  • Room state includes the current stage, permissions, timing, participant list, and facilitator-generated summaries.
  • Room state is rebroadcast in full whenever components change, allowing clients to converge without reconstructing event history.

B.2.1 Event Broadcasting.

Single-participant system messages are broadcast to all clients and filtered locally, while participant laptops may experience transient Wi-Fi disconnections. The communication layer accounts for these conditions through client-side filtering and message identifiers.

  • Single-participant system messages are broadcast to all clients and filtered client-side so only the intended recipient renders them.These messages contain turn-taking instructions but no discussion content.
  • Transient participant-side Wi-Fi disconnections were possible, so the communication layer was designed to tolerate them.The server and orchestration module shared one physical machine, avoiding network issues between those modules.

B.2.2 Fault Tolerance.

The system represents the hierarchical discussion procedure as an explicit state graph and separates event observation from state application. Structured outputs, shared context, and controlled latency further constrain agent behavior and discussion timing.

  • The discussion procedure is implemented as an explicit LangGraph state graph whose nodes represent stages, phases, and generation steps.The graph makes transition rules inspectable and supports isolated execution of subgraphs for testing.
  • An observer queues incoming server events while nodes execute, then applies pending events at node boundaries to prevent mid-call state changes.This separates event observation from application during LLM calls.
  • Structured Outputs constrains speakers, answerer selection, and question references to entities and ideas present in the discussion.The required schemas prevent wrong-speaker attribution, nonexistent participants, and questions about unraised ideas.
  • Agents receive current-stage utterances and selected summaries from preceding stages rather than maintaining private memory.Reviews enter later stages through the same shared summary format.
  • Agent responses are released only after generation and configured pre- and post-utterance delays have elapsed.Internal ChatGPT response times were 800–2,000 ms, while delays for utterances of at least ten words were at least 4,000 ms.

B.4.3 Experimenter Involvement.

The experimenter initiated sessions and stage transitions but did not intervene after the self-thinking phase, leaving facilitation decisions to the facilitator agent and system. The scenario and prompts specified the product context, stage guidance, speech rules, and summary operations used during discussions.

  • Experimenter Involvement: After initiating each session and measuring the two-minute reflection period, the experimenter did not intervene further.All subsequent facilitation decisions were produced by the facilitator agent and system.
  • Scenario Prompt: The discussion scenario asked participants to act as engineers designing Memora, a memory-augmentation smart contact lens with recording, streaming, and AI-assisted memory features.The scenario was based on a Black Mirror design-fiction episode and provided a consistent topic for randomly matched groups.
  • Facilitator Prompts: Facilitator prompts explain the current stage and scenario, announce timed ideation, and encourage participation or stage-focused discussion when needed.Common speech rules require concise, human-readable contributions restricted to the current discussion stage.
  • Summarization Prompts: Dialogue summarization prompts add or modify ideas when needed and finish when content is unrelated or already appropriately summarized.The stakeholder stage additionally classifies items as direct, indirect, or excluded stakeholders.
Loading 2609.11529v1…