Source-linked AI summary
Pack It My Way: Triadic Human-Robot Collaboration for Personalized Autonomous Packing
Sandeep Chowdary Kotapati, Yanxin Gao, Tsung-Chi Lin
TL;DR
Personalized autonomous packing must account for resident preferences that scene geometry alone cannot determine, while continuous expert mediation limits scalability. This paper compares human-expert and voice-agent mediation in a triadic user study using Show–Correct–Generalize across three preference categories. Voice-agent outcomes were comparable in two categories despite shorter instructions, while perceived reliability and preference generalization remained challenges.
Problem
Personalized packing preferences are not fully inferable from scene geometry, yet continuous expert mediation limits scalable autonomous assistance.
Method
A within-subjects study compared human-expert and voice-agent correction across Protection, Compactness, and Grouping using Show–Correct–Generalize.
Results
Voice-agent mediation achieved outcomes comparable to human-expert mediation in two of three preference categories despite shorter and less detailed instructions.
Takeaways & Limitations
Voice agents may reduce continuous expert involvement when preferences can be expressed through structured relational commands, but reliability and semantic generalization remain challenges.
Takeaways & Limitations
The study used a small university sample, simplified objects, marker-based perception, constrained voice commands, one expert, and one-shot corrections.
Abstract
from arXiv · showhide
Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred from scene geometry alone. Expert teleoperators can interpret these preferences and translate them into feasible robot actions, but continuous expert involvement limits scalable deployment. In this paper, we investigate triadic human-robot collaboration among a resident, a correction mediator, and a robot by comparing human-expert and voice-agent mediation. We evaluate the two conditions in a user study across Protection, Compactness, andGrouping tasks, using a Show-Correct-Generalize process to assess preference correction and subsequent generalization after the surrounding objects are rearranged. Results show that voice-agent mediation achieves outcomes comparable to human-expert mediation in two of the three preference categories, despite receiving shorter and less detailed instructions. Both mediators are similarly easy to use, although the human expert is perceived as more reliable. These findings demonstrate the potential of voice agents to reduce expert involvement while identifying perceived reliability and preference generalization as remaining challenges.
I. INTRODUCTION
Personalized autonomous packing requires mediation between resident preferences and robot actions because those preferences are not fully inferable from scene geometry. This paper compares human-expert and voice-agent mediation in a triadic collaboration framework and evaluates correction, generalization, communication, and user perceptions.
- Motivation: The triadic framework assigns preference ownership to the resident, correction mediation to a human expert or voice agent, and physical execution to an autonomous robot.The mediator translates the resident’s desired outcome into an executable correction.
- Motivation: Continuous human-expert involvement can interpret situated preferences but limits scalable personalized assistance and progression toward autonomy.Voice agents are proposed as an alternative that translates spoken preferences into robot actions.
- Study Design: The study compares human-expert and voice-agent mediation across Protection, Compactness, and Grouping preference categories.Protection concerns safety and constraints, Compactness concerns spatial organization, and Grouping concerns semantic relations.
- Study Design: Each task uses a Show–Correct–Generalize process in which the robot demonstrates behavior, receives a resident-specific correction, and reapplies the preference after rearrangement.This process evaluates both correction and subsequent preference generalization.
- Findings: Voice-agent mediation achieved outcomes comparable to human-expert mediation in two of three preference categories despite shorter and less detailed instructions.The mediators were similarly easy to use, but the human expert was perceived as more reliable.
- Contributions: The paper contributes a triadic collaboration framework, the Show–Correct–Generalize process, and a user-study comparison of mediation methods.The evaluation also examines communication efficiency, preference generalization, and user perceptions.
II. BACKGROUND AND RELATED WORK
Prior work advances autonomous packing, preference learning, and language-based feedback, but provides limited support for situated resident preferences and for understanding how mediation affects correction and generalization.
- Autonomous Packing: Existing packing systems generally optimize geometric, physical, or semantic objectives rather than resident-specific preferences that change across contexts.Examples include packing density, stability, fixed objectives, aggregated demonstrations, and model-derived constraints.
- Research Gap: The paper addresses these gaps by studying resident-specific packing preferences and comparing how different mediators affect correction and subsequent behavior.The comparison is situated within a triadic human–robot collaboration setting.
- Learning from Human Feedback: Human-feedback methods let people refine robot behavior through demonstrations, corrections, evaluations, and preference comparisons.Natural language has also been used as an accessible feedback channel in shared-autonomy systems.
- Research Gap: Most prior feedback studies model a direct teacher–robot interaction rather than a separate mediator translating user preferences into executable corrections.They therefore do not isolate how mediation changes communication, execution, trust, or generalization.
C. Triadic Human-Robot Collaboration
Triadic collaboration extends human–robot interaction by assigning distinct roles to a preference owner, a correction mediator, and an embodied executor. The paper focuses on whether an automated voice agent can mediate resident preferences with outcomes comparable to human expertise.
- Prior Triadic HRI: Prior multi-agent HRI systems involve multiple humans, robots, or computational agents with distinct roles, communication channels, and capabilities.Triadic frameworks have supported assistive tasks involving complementary physical capabilities and objectives.
- Research Gap: Earlier triadic systems largely assume fixed team compositions and prescribed roles rather than preference mediation.Preference mediation requires one person to own the preference, another agent to translate it, and a robot to execute and generalize it.
- Research Gap: Three gaps motivate the study: limited attention to situated packing preferences, limited comparative evidence for mediation, and limited knowledge about post-rearrangement generalization.The preference categories include Protection, Compactness, and Grouping.
- Study Scope: The paper compares human-expert-mediated and voice-agent-mediated correction across three preference categories using a Show–Correct–Generalize process.The comparison examines communication, correction quality, reliability, satisfaction, and autonomous generalization.
III. TRIADIC HUMAN-ROBOT COLLABORATION
The proposed collaboration separates preference ownership, correction mediation, and physical execution while preserving a common autonomous manipulation pipeline. Residents communicate desired outcomes, mediators translate them into feasible corrections, and the robot learns relational placement goals for later generalization.
- Preference Owner: The preference owner determines acceptable outcomes using contextual and personal knowledge unavailable to the mediator or robot.The resident communicates the desired packing outcome rather than robot trajectories or low-level manipulation commands.
- Preference Owner: The paper considers safety and constraint, spatial-organization, semantic, aesthetic, and privacy or boundary preferences.Examples include fragility sensitivity, packing density, grouping by object type, neatness, and restricted areas.
- Embodied Executor: The embodied executor uses a 7-DoF Kinova Gen3 arm, a Robotiq 2F-85 gripper, calibrated cameras, and interfaces supporting monitoring and teleoperation.The perception system detects object markers and resident hand gestures while cameras support workspace observation and visual servoing.
- Embodied Executor: The robot autonomously estimates object pose and transports objects to specified targets through Cartesian-space pose control.Preference learning updates the relational placement goal while retaining the underlying manipulation pipeline.
- Correction and Generalization: The Show–Correct–Generalize framework separates autonomous pickup and transport from preference-based positioning mediated by an expert or voice agent.The robot demonstrates an initial placement, receives a correction, and generalizes the learned relation after reference objects are rearranged.
C. Correction Mediator
The correction mediator bridges resident preferences and robot execution through either expert teleoperation or voice-based commands. Both approaches convert corrections into relational placement goals that support autonomous execution and preference generalization.
- Human-Expert Mediation: Human-expert mediation combines workspace observation, resident guidance, and teleoperation to produce a resident-specific placement correction.The expert interprets guidance with respect to reachability, gripper geometry, and workspace constraints before releasing the object.
- Voice-Agent Mediation: Voice-agent mediation transcribes a single spoken instruction and maps constrained vocabulary to predefined robot actions.The system uses local Whisper transcription and a keyword-to-action mapping rather than a general-purpose LLM.
- Command Translation: Supported voice commands include basic qualitative relations and parameterized forms with reference objects, directions, and metric distances.The command structures span the three resident preference categories.
- Command Translation: A deterministic parser converts each transcript into a relational command containing a reference object, spatial relation, and optional distance.Qualitative terms such as tightly and loosely map to predefined distances before conversion into a Cartesian placement target.
- Shared Representation: Both mediation approaches update a resident-specific relational placement goal while leaving the robot’s basic manipulation pipeline unchanged.After surrounding objects move, the stored relationship is applied to updated locations to compute a new placement target.
IV. USER STUDY
The user study used a within-subjects design to examine how human-expert and voice-agent mediation affect preference communication, packing quality, reliability, and autonomous generalization. Twelve university participants completed the study, with sessions lasting approximately 90 minutes.
- Study Design: The study compared human-expert and voice-agent mediation in a within-subjects user study.The evaluation targeted preference communication, packing quality, perceived reliability, and subsequent autonomous generalization.
- Participants: 12 participants from a local university assumed the resident role in the study.Participants included 8 men and 4 women, aged 19 to 40 years.
- Procedure: Each study session lasted approximately 90 minutes, and participants received $15 in compensation.
B. Preference-Based Packing Tasks
The study defines three packing tasks that isolate safety and constraint, spatial organization, and semantic organization preferences. In each task, participants specify Object 3’s placement relative to two reference objects, whose shared geometry and mass control manipulation difficulty.
- Task Categories: Three packing tasks isolate safety and constraint, spatial organization, and semantic organization preferences.The tasks are Protection, Compactness, and Grouping, respectively.
- Task Setup: Across tasks, Object 3 is placed relative to two reference objects while object geometry and mass remain constant.Color and ArUco markers distinguish the objects and allow preference context to vary without changing manipulation difficulty.
- Protection: In the Protection task, participants specify how the soft target should protect the fragile object from the heavy object.Object 1 is heavy, Object 2 is fragile, and Object 3 is soft.
- Compactness: In the Compactness task, participants specify how tightly or loosely Object 3 should be packed relative to two reference objects.
- Grouping: In the Grouping task, participants place Object 3 near the reference object sharing its assigned semantic category.
C. Show–Correct–Generalize Framework
The Show–Correct–Generalize framework evaluates whether a robot can acquire a resident-specific relational preference from one correction and apply it after the scene changes. The study compares human-expert and voice-agent mediation across randomized, manually reset executions.
- Show: Show establishes a preference-agnostic baseline by autonomously placing Object 3 before preference information is provided.This lets the resident observe the initial behavior and identify the need for correction.
- Correct: Correct captures the resident’s preference under the same object configuration through either the human expert or voice agent.The correction is encoded as a relational placement goal defined with respect to Objects 1 and 2.
- Generalize: Generalize rearranges the reference objects and tests whether the robot transfers the stored relationship rather than reproducing an absolute position.The robot applies the relationship to updated reference locations and autonomously places Object 3.
- Procedure: The experiment compared both mediators across three packing tasks in a within-subjects design with randomized task order.Scenes were manually reset before each execution, and the Show execution was independent of the mediator.
- Mediator Conditions: Voice-agent trials used reference-card command formats, whereas human-expert trials allowed unconstrained continuous verbal guidance.The same expert conducted all expert-mediated trials and had more than 200 hours of teleoperation experience.
- Evaluation: Participants evaluated each correction process, each generalization outcome, and the overall comparison of the two mediators.
E. Measures and Analyses
The study measures correction communication, task-specific packing quality, and participants’ perceptions of mediator usability and reliability. Packing quality is analyzed across correction and generalization, with mediator comparisons using equivalence and paired tests.
- Correction Communication: Instruction complexity was classified as Brief, Moderate, or Detailed using word counts and action or refinement criteria.Brief means no more than 10 words and one relation or action; Detailed means more than 25 words, repeated refinements, or at least four actions.
- Packing Quality: Task-specific packing quality was normalized to [0, 1] using geometric criteria tailored to Protection, Compactness, and Grouping.The measures use object-center positions, perpendicular deviation, triangle area, and related-versus-unrelated object distances.
- Packing Quality: Higher packing-quality values indicate stronger satisfaction of the corresponding task criterion.Thresholds were anchored to physical object dimensions, including half the object width and the object footprint area.
- Analysis: The analysis compared mediator conditions with paired-samples TOST using equivalence bounds of ±0.10 on the normalized 0–1 scale.Other outcomes used paired-samples t-tests or Wilcoxon tests depending on within-participant difference distributions.
A. Correction Communication
Voice-agent mediation used shorter instructions while matching human-expert correction-stage packing quality for Compactness and Grouping, but not Protection. Generalization depended on preference type, and perceived reliability remained a challenge despite similar ease of use.
- A. Correction Communication: Voice-agent instructions were shorter for Protection and Grouping, while Compactness showed the same nonsignificant tendency.Voice-agent instructions were predominantly Brief; human-expert corrections contained more Moderate and Detailed instructions.
- A. Correction Communication: Correction-stage packing quality was equivalent between mediators for Compactness and Grouping, but equivalence was not demonstrated for Protection.Paired TOSTs supported equivalence for Compactness and Grouping with both p < .05.
- A. Correction Communication: Shorter voice-agent instructions and equivalent quality for Compactness and Grouping suggest reduced expert involvement for structured preferences supported by the command space.The paper proposes retaining human experts for ambiguous, unsupported, or higher-stakes preferences.
- B. Task-Dependent Preference Generalization: Compactness quality increased during Generalize under both mediators, while Grouping decreased under voice-agent mediation and Protection showed no significant change.The Compactness increase reflects a higher-quality rearrangement from the stored relational goal rather than additional learning.
- B. Task-Dependent Preference Generalization: Geometric preferences transferred similarly across mediators, but fixed spatial representations may not preserve semantic Grouping preferences after scene changes.The paper recommends explicit semantic relationships, clarification, and target confirmation for future voice mediators.
- C. Ease of Expression and Perceived Reliability: Eight participants found the mediators equally easy to use, whereas five considered the human expert more reliable.Voice-agent mediation nevertheless produced greater learning satisfaction for Protection, with no mediator difference for the other tasks.
- C. Ease of Expression and Perceived Reliability: Ease of expression did not necessarily translate into perceived reliability, motivating visible interpretations and resident confirmation before execution.The paper connects this trust gap to possible human-expert advantages in interpreting ambiguity and recovering from incorrect guidance.
VI. CONCLUSION
The framework compares human-expert and voice-agent mediation in triadic personalized packing using Show–Correct–Generalize. Voice agents can reduce expert involvement, but generalization and study scope remain constrained.
- Voice-agent mediation produced shorter, less detailed instructions and achieved outcomes comparable to human-expert mediation in two preference categories.The comparison spans Protection, Compactness, and Grouping tasks.
- Protection quality did not significantly decline after rearrangement under either mediator, while Compactness quality increased after rearrangement.Figure 6 compares Correct and Generalize packing-quality scores within each mediator and task.
- Grouping quality declined under voice-agent mediation after rearrangement, suggesting that predefined spatial representations may not fully preserve semantic preferences.The conclusion identifies semantic preference preservation as a remaining challenge.
- Most participants rated both mediators equally easy to use, although more participants considered the human expert more reliable.The overall preference panel summarizes ease of expression and reliability.
- The study was limited by a small university sample, simplified objects, marker-based perception, constrained voice commands, and one human expert.It also examined one-shot corrections rather than preferences learned through repeated household interactions.