Source-linked AI summary
Teaching is a Process: The TOSS Framework for Modeling Human Teaching Decisions in Human-Interactive Robot Learning
Bernhard Hilpert, Kim Baraka, Joost Broekens
TL;DR
Current interactive teaching settings can misalign teachers’ intuitive procedures with robot learning mechanics, prompting strategic adaptation. This work presents TOSS, a procedural framework in which teaching decisions connect triggers, objectives, signals, and strategies, supporting more sophisticated computational teacher models.
Problem
Current HIRL settings often place teachers in reactive roles and may produce teaching behavior shaped by misalignment with the robot’s inferred learning mechanics.
Method
The TOSS framework models teaching as an iterative process linking perceived triggers, subjective objectives, communicative signals, and high-level strategies.
Results
Teaching decisions are modeled as an integrated, multifaceted procedural loop in which mentors align subjective targets and interpretations of robot shortcomings with appropriate teaching signals.
Takeaways & Limitations
TOSS provides a structured basis for computational teacher models, realistic teaching simulations, and interaction strategies that respond to shifts in teacher strategy.
Takeaways & Limitations
The framework’s trigger dimensions may overlap, and objective categories have fluid boundaries that may not map directly onto technical feature spaces.
Abstract
from arXiv · showhide
Successful Human-Robot Teaching assumes alignment between robot processing needs and human teaching intent. To better understand this alignment, this work seeks to uncover the underlying logic that humans intuitively apply when teaching. Through an exploratory, bottom-up study with N=34, participants observing two distinct robot Reinforcement Learning (RL) scenarios, we analyze 204 intuitive teaching responses across early, middle, and late learning phases. Results reveal that teaching decisions consist of a nuanced, interconnected network of Triggers (situational catalysts), Objectives (subjective teaching targets), Signals (communicative acts), and Strategies (high-level governance) in which teachers spontaneously adopt diverse roles, acting as coaches, engineers, or designers and prioritize different objectives. Based on these results, we introduce the TOSS Framework, which conceptualizes Human-Robot teaching as a procedural loop between robot behavior and human teaching actions, in which human teaching decisions are modeled as Trigger-Signal responses modulated by teaching Objectives and Strategies. It provides future research with an openly accessible dataset and a theoretical foundation for a) understanding teaching decisions and b) simulating realistic oracles as well as c) designing human-centered teaching settings and novel robot learning algorithms that go beyond the constraints of current robot learning settings.
I. INTRODUCTION
The paper argues that current human-interactive robot learning studies mainly capture reactive teaching adaptations, leaving intuitive teaching priorities unclear. It addresses this gap by decoupling teachers from immediate interaction pressures to study their baseline teaching decisions.
- Research gap: Current HIRL paradigms often place teachers in reactive roles that may produce suboptimal teaching behavior.Prior work has examined post-hoc justifications, mental models, and algorithmically inferred goals, but interactive settings can distort teachers’ preferred strategies.
- Approach: The work uses an unconstrained, noninteractive experiment across three learning intervals and two RL settings to describe raw teaching triggers, objectives, signals, and strategies.The settings involve tabular Q-learning for navigation and a DDPG agent for manipulation.
- Research gap: Interactive teaching settings can force teachers to compensate for mismatches between their intuitive teaching policy and robot learning mechanics.Such compensation may cause teachers to adapt or “hack” their signals rather than express their preferred procedural approach.
- Research aim: The study examines what teachers intuitively prioritize when they are not forced into strategic compensation.Its stated goal is to uncover a baseline teaching process by decoupling teachers from immediate interaction pressures and loop dependencies.
III. METHODOLOGY
The methodology used an open, qualitative observational study in which participants watched robots learn without intervening. Two robot embodiments and task types were used to examine teaching intuitions across different learning settings.
- Study design: The experiment was preregistered, IRB-approved, and made its questionnaire, coding plan, stimuli, data corpus, and protocols openly available.These materials were provided through OSF.
- Participants: N = 34 participants remained after one inattentive respondent was excluded from the recruited sample.Participants were recruited through Prolific and ranged from 19 to 65 years old.
- Procedure: Participants observed an RL robot learning process while framed as teachers, but they could not intervene during the study.The questionnaire presented two robot-learning scenarios after briefing and consent procedures.
- Experimental scope: The experiment used two embodiments and tasks to test whether teaching intuitions extend across tabular and function-approximation learning.One robot learned grid-world navigation and the other learned manipulation with a robotic arm.
- Experimental scope: The study therefore compared teaching observations across distinct robot learning settings rather than a single task or learner type.This design was intended to assess whether teaching intuitions apply across the two learning paradigms.
3) Scenarios:
The scenarios showed a cleaning robot learning navigation and an assistive robotic arm learning medication manipulation. Training trajectories were rendered at regular checkpoints to make learning progression observable.
- Manipulation task: The manipulation scenario used a DDPG agent learning to push medication toward a target with a robotic arm.The figure identifies medication as blue and the target as yellow.
- Stimulus construction: Learning progression was represented by stimulus videos generated from trajectories sampled at regular training intervals.Convergence was defined as five consecutive successful task completions.
- Stimulus construction: Navigation checkpoints were saved every 500 episodes during 10,000 episodes, while manipulation checkpoints were saved every 1,000 steps during 100,000 steps.The resulting videos lasted 94 seconds for navigation and 192 seconds for manipulation.
- Analysis: The analysis used inductive, reflexive thematic analysis with data reduction, theme generation, and refinement against the full dataset.The pipeline condensed statements, clustered latent themes, and cross-validated them for internal consistency.
IV. RESULTS
Participants’ statements portrayed teaching as a heterogeneous, multi-faceted process rather than a single act. The results organize this process into Triggers, Objectives, Signals, and Strategies, with Triggers classifying the situations that prompt intervention.
- Overall findings: Teachers produced heterogeneous statements and intuitively conceptualized teaching as a nuanced, multi-faceted process.This pattern appeared when participants were not instructed how or when to intervene.
- Overall findings: The framework distinguishes Triggers, Objectives, Signals, and Strategies as four facets of intuitive teaching decisions.They represent situational catalysts, subjective targets, communicative acts, and high-level governance, respectively.
- Triggers: Triggers are observed situated robot behaviors that prompt a teaching approach, reflecting conflict between a teacher’s mental model and observed behavior.They are not defined as objective robot states.
- Triggers: Consistency distinguishes systematic perceived behavior patterns from incidental discrete events or glitches.For example, repeated overshooting is systematic, whereas a dramatic horizontal movement may be incidental.
- Triggers: Reference frame distinguishes relative evaluations against the robot’s observed progress from absolute evaluations against task logic.A robot can be technically suboptimal yet relatively improved, or violate task requirements regardless of prior behavior.
- Triggers: Alignment distinguishes desired behaviors that fit teaching goals from undesired behaviors that violate them.Undesired triggers identify shortcomings requiring corrective or preventive intervention.
- Triggers: Trigger dimensions intersect into clusters that represent how teachers categorize robot behavior before intervening.These clusters are illustrated through thematic anchors defining the dimensions’ boundaries.
B. Objectives: The “What" of Teaching
Objectives are subjective target states describing what teachers want robot interventions to achieve. They span behavioral execution, task outcomes, and the robot’s internal knowledge or understanding.
- Objectives are subjective target states representing the specific goal behaviors teachers intend to achieve through intervention.
- Behavioral Objectives target how the robot executes a task, independently of whether it achieves the task outcome.They include systematism, precision, efficiency, and stability.
- Goal/Task Objectives target milestones and task outcomes independently of execution quality.
- Knowledge Objectives target the robot’s internal model or understanding of the environment and task.Examples include mapping the environment, locating objects, and clarifying object identities or goals.
C. Signals: The “How" of Teaching
Signals are the interactive methods teachers use to communicate how the robot should learn. The framework identifies evaluative feedback, corrective advice, demonstration, information sharing, and withholding.
- Signals define the actual intervention by specifying how the robot should be taught.Five intuitive signal clusters emerged from the data.
- Evaluative Feedback validates or invalidates past or anticipated behavior through positive or negative signals.The teacher acts as a Judge using reinforcement or punishment.
- Corrective Advice gives relative adjustments to past or anticipated wrongful behavior.The teacher acts as a Tutor by specifying a modification to the robot’s attempt.
- Demonstration provides absolute reference trajectories that guide the robot toward target behavior.The teacher acts as an Example for the robot to learn from.
- Information Sharing communicates declarative information about the environment, task, or goal specifications.The teacher acts as an Oracle by supplying contextual or goal information.
- Withholding deliberately refrains from interfering as a functional null signal.The teacher acts as an Observer to foster independence or test learning convergence.
D. Strategies: The Teaching Approach
Strategies govern the teaching process at a high level, defining where teachers intervene and the role or plan they adopt. The framework identifies five distinct strategies, including interactive guidance, system adaptation, task adaptation, laissez-faire, and satisfied supervision.
- Strategies define the teacher’s intervention locus and strategic profile, including the teaching mode or plan for robot development.
- Interactive Guidance uses real-time feedback cycles in which a Coach reacts to Triggers and combines Objectives with Signals.
- System Adaptation treats the robot’s configuration as insufficient and positions the teacher as an Engineer who changes architecture, tools, sensors, or learning algorithms.
- Task Adaptation positions the teacher as a Designer who scaffolds the environment or task complexity to maintain a zone of proximal development.
- Laissez-faire deliberately removes teacher intervention to support autonomous learning, whereas Satisfied Supervision responds specifically to perceived mastery.
E. The TOSS Framework: Teaching as an integrated process
The TOSS Framework models teaching as an integrated process linking perceived robot behavior, subjective teaching targets, communicative acts, and high-level strategy. This process is iterative and structurally flexible rather than a fixed sequence.
- TOSS formalizes teaching as an integrated process of Triggers, Objectives, Signals, and Strategies.
- A complete teaching cycle links a Trigger indicating why intervention occurs, an Objective specifying what to achieve, and a Signal expressing how, under a Strategy.
- Teaching decisions can bypass facets: proactive guidance may omit a Trigger, while other strategies can act without a failure or communicative Signal.
- TOSS portrays teachers as dynamically shifting among Coach, Engineer, and Designer profiles while aligning robot-behavior evaluations with subjective targets.
V. DISCUSSION
The discussion presents TOSS as a framework for the pedagogical logic underlying human-robot teaching decisions. It emphasizes proactive evaluation and fine-grained, subjective objectives beyond simple task-failure responses.
- TOSS complements algorithm-focused HIRL research by modeling the pedagogical logic of teaching decisions.
- Triggers arise from proactive internal evaluation of robot behavior rather than merely reacting to discrete behavioral errors.
- Teachers distinguish execution-level and goal-oriented objectives, revealing pedagogical targets finer-grained than high-level reward functions.
C. Signals
TOSS treats teaching signals as connectors between diagnosed robot behavior and subjective teaching targets, within a flexible strategy-dependent process. This expands teaching beyond isolated feedback acts or a rigid linear sequence.
- Teaching signals include evaluative feedback, corrective guidance, demonstrations, information sharing, and withholding, but are not isolated communicative acts.
- Signals connect a diagnosed Trigger to a subjective Objective, such as pairing aimless wandering with goal-setting information or demonstration.
- Strategies make teaching decisions dependent on the teacher’s perceived strategic identity and support dynamic iteration rather than a rigid T-O-S sequence.
- The framework links teaching profiles that earlier work mostly studied separately, supporting systems that let teachers shift between teaching and re-architecting the learning task.
- TOSS allows teaching to be modeled as modular, non-linear, and structurally flexible, including preemptive instruction that skips situational Triggers.
F. Future Work & Limitations
TOSS offers a flexible foundation for modeling human teaching while remaining an exploratory baseline whose categories require broader validation and refinement. Future work should test the framework quantitatively and in closed-loop studies, while designing more continuous, multi-level teaching interactions.
- Limitations: The TOSS categories are fluid: Trigger dimensions may overlap, and Objective boundaries can intersect across functional levels.The framework should disambiguate subjective teaching targets without treating them as rigid, mutually exclusive categories.
- Limitations: The findings provide an exploratory glimpse into human pedagogical logic rather than a definitive taxonomy.The authors characterize the results as a foundational baseline requiring further study.
- Future Work: Large-scale quantitative validation and closed-loop interactive studies are needed to assess teaching behavior with a live, adapting robot.These studies would extend the current high-fidelity but non-definitive characterization of teacher conceptualization.
- Future Work: TOSS can support computational teacher models, realistic oracle simulations, and robot interaction strategies that respond to shifts between Coach and Engineer roles.Combining TOSS’s subjective perspective with quantifiable feedback models could move research beyond simple signal matching and help mitigate mental-model mismatches.
- Future Work: The framework’s structural flexibility motivates teaching settings with continuous, multi-level feedback rather than rigid turn-taking interfaces.The proposed direction emphasizes systems co-developed between teachers and robots.
- Future Work: TOSS contributes an openly accessible dataset and theoretical foundation for understanding teaching decisions and designing human-centered robot learning systems.The stated scope includes realistic oracles, human-centered teaching settings, and novel Human-Interactive Robot Learning algorithms.