Source-linked AI summary

Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion

Yue Shen, Rehema Abulikemu, Ryan P. McMahan, Yan Chen

arXiv:2608.26185v1cs.AIcs.ETcs.HC

TL;DR

People may withhold useful points in co-located discussion because speaking exposes them to interpersonal or professional risks. SecondVoice uses a mixed-reality embodied proxy and structured intent specification to bring such points into the spoken exchange; in a preliminary N = 16 comparison, participants reported using it for withheld points and proxy deliveries received multi-turn engagement, with tradeoffs around timing, ownership, and trust.

  • Problem

    People may withhold concerns, disagreement, or suggestions in co-located discussion when they anticipate negative social or professional judgments.

  • Method

    SecondVoice uses a shared embodied proxy and private structured specification so users can specify intent without composing a full utterance.

  • Results

    Three observed proxy deliveries were followed by multi-turn, multispeaker exchanges, which were not found after text-board posts.

  • Takeaways & Limitations

    SecondVoice offers a situationally valuable participation channel while introducing tradeoffs around timing, authorship, ownership, and trust in reformulation.

  • Takeaways & Limitations

    The comparison bundled input method, authoring effort, embodiment, output modality, timing, and spoken-floor entry, preventing attribution of observed differences to one factor.

Abstract

from arXiv · show

Equal participation in co-located discussion is important for effective collaboration, yet people often hold back when they anticipate negative interpersonal or professional consequences, especially when raising a point requires voicing it themselves. We present SecondVoice, a mixed-reality system that enables people to speak up through an embodied virtual proxy. By separating what is said from who says it, SecondVoice brings hesitant points into the live spoken discussion without putting the speaker on the spot. Using a private overlay, users specify their intent through a structured specification process rather than composing a full utterance. The system reformulates the input and voices it into the conversation through the proxy. We characterize a design space of participation channels under social risk. In a preliminary within-subject study (N = 16), we compare the complete SecondVoice system with an anonymous text-board channel across two group discussion tasks. Half of participants reported using SecondVoice for a point they did not say aloud, compared with 18.8% for the text board. Proxy-delivered points entered the spoken floor and were followed by multi-turn group engagement, which we did not observe after text-board posts. Participants described the channel as situationally valuable but identified tradeoffs around timing, ownership, and trust in reformulation.

1 Introduction

SecondVoice addresses the reluctance to self-voice socially risky points by using a shared embodied proxy to bring privately specified remarks into live discussion. A preliminary comparison found greater spoken-floor engagement after proxy deliveries than after anonymous text-board posts, alongside tradeoffs in timing, trust, and ownership.

  • System and motivation: People often withhold concerns, disagreement, or suggestions because they fear being judged as uninformed, disruptive, or difficult to work with.This reluctance contributes to unequal participation in classrooms and workplaces despite relevant expertise among quieter participants.
  • Prior approaches: Anonymous text boards can broaden participation but often leave remarks in a side layer, while embodied agents may occupy a more independent social role.SecondVoice explores a bounded middle position: a proxy that voices only what hesitant participants privately specify.
  • System and motivation: SecondVoice lets hesitant participants have privately specified points voiced by a shared embodied proxy in co-located discussion.Users specify intent through a structured process, and the system reformulates and delivers the point through the proxy.
  • Preliminary comparison: Three observed proxy deliveries were followed by multi-turn, multispeaker exchanges, which were not found after text-board posts.Participants also described the proxy as reducing pressure to speak in their own voice, while raising concerns about timing, trust, and ownership.
  • Contributions: The paper contributes a proxy-mediated participation system, a structured intent-specification process, and preliminary comparative findings against an anonymous text channel.The contribution includes a design space characterizing tradeoffs among participation channels under social risk.

2 Related Work

Prior systems broaden participation through parallel channels, facilitation, or message reshaping, but they often leave contributions outside the spoken exchange or alter perceived ownership. SecondVoice investigates a shared embodied proxy that voices participant-confirmed points while preserving a private input layer.

  • Parallel channels: Backchannels and anonymous text boards reduce direct speaking pressure but often keep contributions in a side layer that others may overlook.In co-located settings, visible typing, identifiable handles, and traceable logs can also weaken anonymity.
  • Participation support: Participation displays, sociometric sensing, turn-taking visualizations, and reflective dashboards support more balanced interaction at different stages of discussion.These systems regulate or reflect participation rather than voicing a specific participant’s point into the live exchange.
  • Message mediation: AI-mediated communication reshapes message form through augmenting, rewriting, or generating text, which can affect users’ perceived agency and ownership.This approach changes how a message sounds rather than directly regulating who takes the floor.
  • Embodied agents: Discussion agents act as facilitators, advocates, or independent social actors, but a bounded proxy that voices privately specified participant points remains underexplored.The gap concerns synchronous, co-located discussion where one shared proxy can speak on behalf of present participants.
  • Embodiment: Embodied agents can receive higher rapport and trust ratings and support more balanced turn-taking than disembodied counterparts.Virtual appearance, behavior, and multimodal attention cues also influence engagement, inclusivity, satisfaction, and noticeability of a new speaker.
  • SecondVoice: SecondVoice combines a private participant input layer with a shared co-present proxy, and each delivery voices one participant-confirmed point.This differs from summarizers that aggregate views and from a shared display alone.

3 Participation Channels Under Social Risk

Participation channels trade off direct exposure, expressive specificity, and entry into the spoken floor. The paper identifies an underserved channel that lowers exposure while preserving nuanced expression and live spoken-floor entry, using a bounded proxy to mediate delivery.

  • Design dimensions: Participation channels differ in direct exposure, expressive specificity, and whether a point enters the shared spoken floor.The spoken floor is the group conversation, where a point enters as a turn others can answer.
  • Existing channels: Direct speech offers high expressive specificity and direct spoken-floor entry but ties the point to public self-voicing in the moment.Anonymous text boards and backchannels lower exposure while typically keeping remarks in a side layer; reactions and polls reduce burden but express less.
  • Underserved region: The underserved region combines nuanced expression, lower direct exposure, and entry into the live spoken discussion.Existing channels generally provide either lower exposure outside the spoken floor or spoken-floor entry at the cost of direct exposure.
  • Proxy-mediated delivery: A proxy can partially separate a point’s content from the author’s exposure by allowing someone other than the author to voice it.This separation addresses where a point appears and who voices it.
  • Bounded proxy: The underexplored proxy role voices only what a user has specified and confirmed, rather than determining content independently or replacing the user entirely.This bounded role positions the proxy between autonomous agents and full user relays.
  • Timeliness and authorship: Delegated delivery creates a timeliness–authorship tradeoff: autonomous generation keeps pace, while full-utterance relays preserve authorship but may arrive too late.SecondVoice’s structured specification offers a guided alternative to composing a complete utterance.

4 SecondVoice

SecondVoice supports participation through three stages: users specify an intended point privately, the system reformulates it in context, and an embodied proxy delivers it on the spoken floor.

  • System Overview: SecondVoice organizes proxy-mediated participation into specification, contextual reformulation, and proxy delivery.Discussion grounding runs across all three stages to support candidate generation, reformulation, and delivery timing.
  • Participation Entry and Specification: Private pulses and nudges offer optional entry points based on participation state, unresolved pins, and recent discussion activity.The system prioritizes pin reminders, then context-specific pulses, then generic nudges; both mechanisms are private and dismissible.
  • Participation Entry and Specification: Users select a communicative move, anchor it to an active discussion topic, and browse candidate points instead of composing a full utterance.The interface offers question, challenge, agree, and suggest moves, with refreshable alternatives for the selected move–topic pair.
  • Contextual Reformulation: After confirmation, the system reformulates the point into context-appropriate speech while preserving the user’s intended stance without adding unsupported elaboration.Reformulation adapts tone and phrasing to the live discussion.
  • Proxy Delivery: The proxy raises its hand and delivers the reformulated point in its own voice without naming the contributor or presenting it as a relayed message.The proxy therefore enters the spoken discussion as a co-present participant rather than a flagged text channel.
  • Proxy Delivery: Delivery uses a queue that checks timeliness, topic shifts, and duplicate points before waiting for a brief conversational pause.Waiting for a pause separates interface operation from delivery but can make a point less timely.

5 User Study

The within-subject study compared SecondVoice with an anonymous text board across two co-located group discussion tasks, using logs, transcripts, observations, questionnaires, and interviews to examine participation and uptake.

  • Study Design: The study compared SecondVoice with an anonymous text board because both let participants contribute without directly voicing a point, while only proxy delivery placed it on the spoken floor.Both conditions used MR headsets and private visual interaction layers.
  • Participants and Procedure: Sixteen participants formed four groups of four, balanced by self-described vocality, and each group experienced both conditions across both tasks.Condition order and task assignment were fully counterbalanced across groups.
  • Tasks: The University President task used shared and private candidate information, while Lost at Sea required consensus ranking of 15 survival items.The tasks created opportunities for unpopular information, disagreement, critique, and minority positions.
  • Participants and Procedure: Participants trained on both channels, completed 10–15-minute group discussions in each condition, then completed surveys and post-study interviews.Training lasted about five minutes per channel, and a 10-minute break separated the discussion rounds.
  • Analysis: Analysis combined backend event logs, audio transcripts, observed spoken uptake, 7-point questionnaires, and thematic analysis of all 16 interviews.Observed uptake tracked whether points were noticed, answered, and carried forward in spoken discussion.

6 Results

SecondVoice was used selectively to voice hesitant points and bring some of them into spoken group discussion, while participants reported tradeoffs involving timing, ownership, authorship inference, and reformulation. Compared with the anonymous text board, proxy-delivered points received more direct conversational uptake, although the study’s quantitative and task outcomes were descriptive.

  • Proxy Use and Participation Patterns: 13 proxy deliveries and 6 text-board posts were logged, with SecondVoice used by participants across both halves of the session speaking distribution.10 of 19 total channel-use events came from participants in the lower half of their group’s speaking distribution, including 5 proxy deliveries.
  • Proxy Use and Participation Patterns: 8 of 16 participants reported using SecondVoice for something they did not say aloud, compared with 3 of 16 for the anonymous text board.The paired difference was not statistically significant and was reported as a descriptive trend.
  • Group Uptake: 3 of 13 proxy deliveries developed into multi-turn, multispeaker exchanges, whereas no comparable spoken sequence followed the 6 text-board posts.Four proxy deliveries received substantively related human turns within 20 seconds.
  • Group Uptake: Participants found board posts easier to overlook because they remained in a parallel layer, while proxy speech claimed a turn and reorganized conversational flow.Participants described the board as something they could acknowledge briefly and move past.
  • Participant Experience and Tradeoffs: SecondVoice lowered direct attribution pressure, but participants identified costs involving credit, authorship inference, and imperfect reformulation.Five participants described reduced pressure or attribution; authorship could still be inferred from behavior and context, and some candidate suggestions did not match intent.
  • Participant Experience and Tradeoffs: Late delivery could erase the proxy channel’s advantage, and participants viewed SecondVoice as valuable mainly when voicing a point carried social risk.Participants described the channel as situational rather than general-purpose, with timing constrained by the live discussion.

7 Discussion and Future Work

The discussion identifies spoken-floor entry as SecondVoice’s defining distinction, while highlighting tradeoffs involving ownership, expressive fidelity, trust, accountability, and study scope. It also outlines future work in higher-stakes settings, longer-term use, and controlled component comparisons.

  • Spoken Floor Entry: Spoken-floor entry distinguishes the proxy from the text board: the proxy captured attention and supported continued discussion, while board posts remained parallel and easier to overlook.Both channels reduced direct exposure and supported expressive specificity, but only the proxy entered the shared spoken floor.
  • Spoken Floor Entry: The proxy performs the socially costly act of claiming the floor, redirecting attention, and creating space for points users may hesitate to voice directly.The discussion frames interruption as face-threatening and treats floor entry as a design concern alongside content and attribution.
  • Design Tensions: Separating authorship from public voicing creates an ownership tradeoff: reduced exposure can weaken credit and accountability, although proxy delivery may scaffold later direct speech.The paper treats this tension as inherent to partial separation rather than a problem that can simply be designed away.
  • Structured Specification as Interaction Design: Structured specification trades open-ended expressive fidelity for discretion and speed, but candidate mismatch can prevent users from expressing their intended point.The paper proposes greater candidate diversity, user-guided steering, and faster refinement loops.
  • Divergent Agent Perceptions: Interpretations of the proxy ranged from human to translator to counselor, and dismissing it could feel personal because it spoke for a real participant.The proxy functioned as a social object whose meaning was negotiated by the group, not merely as a neutral conduit.
  • Design Tensions: Trust remains fragile because contextual reformulation can misstate tone or intent, and the prototype lacked public correction or retraction after delivery.Users confirmed candidate points before reformulation, but one delivery expanded a confirmed point into a claim judged only partially correct.
  • Study Scope: The preliminary study cannot isolate the cause of observed differences because the conditions bundled input method, authoring effort, embodiment, output modality, timing, and spoken-floor entry.The sample included 16 university students in four groups, and the tasks elicited only moderate social risk.
  • Future Work: Future work should test SecondVoice in higher-stakes settings, examine longer-term adoption and perception changes, and vary bundled channel components separately.Suggested variants include alternative proxy embodiments, audio-only or text-to-speech modalities, and different timing and turn-taking designs.

8 Conclusion

SecondVoice explores whether a bounded proxy can bring hesitant contributions into co-located spoken discussion without requiring direct self-voicing. The preliminary comparison shows that proxy delivery entered the spoken floor while preserving unresolved tensions around expression, pace, and authorship.

  • 8 Conclusion: SecondVoice separates point authorship from public voicing through a bounded proxy that enters the spoken floor for hesitant participants.The system uses structured specification to balance expressive contribution, spoken-floor entry, and reduced direct exposure.
  • 8 Conclusion: Half of participants reported using SecondVoice for a point they did not say aloud, while board contributions remained in a parallel text layer.The comparison was preliminary and involved the complete SecondVoice and anonymous text-board channels.

A Task Materials

The study used two counterbalanced group decision-making tasks, with each group completing both tasks under different participation-channel conditions.

  • A Task Materials: Each group completed both group decision-making tasks, with one task per condition and the order counterbalanced.The tasks were adapted from prior work and designed to elicit disagreement and information exchange.

A.1 Lost at Sea (Survival Ranking)

In Lost at Sea, participants individually ranked salvaged items and then worked together to produce a shared survival ranking against an expert solution.

  • A.1 Lost at Sea (Survival Ranking): Participants ranked 15 salvaged items by survival importance before discussing them to produce a shared group ranking.The expert solution provided a baseline for scoring, and individual disagreement created material for group discussion.
  • A.1 Lost at Sea (Survival Ranking): The hidden-profile task supplied shared information about three candidates plus private facts, requiring discussion to select a candidate.Its structure made private-information sharing relevant to reaching the optimal choice.

B LLM Prompt Templates

The prompt templates divide SecondVoice into analysis, structured specification, proxy expression, facilitation, and supporting coordination functions. Together, they extract discussion context, transform participant intent into speech, and manage follow-up interaction while preserving intended meaning and privacy.

  • B LLM Prompt Templates: SecondVoice uses GPT-4o and GPT-4o-mini with structured JSON output and role-specific system prompts.The implementation distinguishes heavy and light model calls across its prompt pipeline.
  • B.1 System Prompts: The base persona frames the system as a tactful AI meeting facilitator embedded in an augmented-reality environment.It instructs the agent to protect private participant identity and hidden source information.
  • B.2 Observation Pass (GPT-4o): The analyst layer extracts speaker-attributed observations, discussion anchors, stance signals, and open issues from transcript windows.Its output includes the current topic, topic-change status, observations, anchors, and open issues.
  • B.7–B.10 Supporting Calls: Supporting calls expand anonymous-board keywords, bridge follow-ups, generate contributor follow-up options, and gate pulse interventions using a substantiveness check.The typed fallback is used in the anonymous text-board condition, while bridge and follow-up calls remain bounded by the original intent.
  • B.3 Participant Context Update (GPT-4o): The participant-context layer updates per-user stance summaries, active concerns, and stance changes from newly extracted observations.The resulting participant structure stores these updates for later system calls.
  • B.4 Specification Screen Generation (GPT-4o-mini): The specification-screen generator offers four dialogue moves—question, challenge, agree, and suggest—and creates discussion issues, default operations, and a preview.It uses participant and discussion context to generate the initial structured interface.
  • B.6 Cluster Adaptation (GPT-4o): Cluster adaptation jointly handles convergent contributions and may drop a point when human speakers have already covered it.The prompt supplies queued contributions, recent transcript context, and delivery history.

C Survey Response Distributions

Figures 8 and 9 show response distributions comparing SecondVoice with the anonymous text board across 13 Likert items and two binary items. No Likert item was significant after correction, while the binary items were analyzed separately.

  • Figures 8 and 9 compare SecondVoice and the anonymous text board across 13 Likert items and two binary items.Figure 8 covers the first eight Likert items; Figure 9 covers the remaining five Likert items and the binary items.
  • No Likert item reached significance after Benjamini–Hochberg correction across the 13-item family.The comparison used responses from N=16 participants.
  • The two binary items were analyzed separately from the Likert family using exact McNemar tests.The binary-item analysis also used N=16 participants.
Loading 2608.26185v1…