Source-linked AI summary
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
TL;DR
Agentic AI lacks a unified paradigm that makes evolving human states and agency the primary objects of modeling, intervention, and evaluation, despite growing personal and embodied applications. The paper introduces Combodied Agents, a closed-loop framework using multimodal perception, longitudinal memory, personal world models, and proportionate intervention policies. It concludes that progress should be judged by sustained human benefit and long-term agency while respecting safety, relationships, and autonomy.
Problem
Existing agents provide fragmented capabilities and often optimize task outcomes without integrating longitudinal human trajectories, agency preservation, and broader wellbeing criteria.
Method
The paper develops a closed-loop Combodied Agent framework with multimodal perception, longitudinal memory, personal world models, intervention policies, feedback, and safety constraints.
Results
The paper proposes a human-centric Agentic AI paradigm, taxonomy, deployment perspective, and evaluation agenda organized around beneficial human trajectories and agency.
Takeaways & Limitations
Combodied Agents treat software, sensors, robots, and human services as action channels while evaluating whether people can understand, choose, correct, recover, and develop capability over time.
Takeaways & Limitations
Hybrid deployments must resolve inconsistent personal-context states, provenance, authority, conflicts, privacy exposure, and safe degradation during connectivity loss.
Abstract
from arXiv · showhide
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.
1 Introduction
Agentic AI has advanced task execution across digital and physical environments, but task success can diverge from human understanding, capability, autonomy, and wellbeing. Combodied Agents reorient modeling, intervention, and evaluation around evolving human trajectories and preserved agency.
- Motivation: Digital and Embodied Agents automate software and physical tasks, while progress is commonly measured by longer horizons, reliable action, and reduced supervision.These systems can reduce costs and automate difficult work, but this framing centers external task completion.
- Motivation: Task performance can diverge from human development when AI improves outputs while reducing understanding, judgment, or cognitive effort.AI-assisted work and decision-making studies report overreliance when users cannot determine when to accept, verify, or reject outputs.
- Proposed direction: Human participation becomes a design problem in which agents choose among independent action, oversight requests, and returning knowledge or capability to the person.Delegation should fit the person’s goals, state, competence, relationships, and risk while preserving control.
- Proposed direction: Combodied Agents use digital tools, wearables, robots, and human services as action channels, with success determined by what people understand, decide, do, and sustain over time.The paradigm complements human-centered AI by treating human trajectories and agency as the primary target.
- Research gap: Personal assistants, health agents, companions, behavioral coaches, and adaptive human–AI systems provide partial capabilities but lack a unified longitudinal human-state paradigm.The missing integration concerns intervention responses, adaptive division of labor, and evaluation of both benefit and preserved agency.
- Proposed direction: Combodied Agents organize perception, memory, prediction, intervention, and evaluation around the evolving person rather than external task or domain outcomes.The framework adds multimodal human-state perception, longitudinal correctable memory, personal world models, calibrated reversible policies, and explicit human-gain evaluation.
2 Foundations of Combodied Agents
The paper introduces Combodied Agents as a formal paradigm and frames the remainder around defining their closed-loop framework and technical components.
- Foundations: The paper presents a canonical definition of Combodied Agents and distinguishes their action substrate from those of Digital and Embodied Agents.It then formalizes the resulting closed loop.
Definition 1: Combodied Agents
Combodied Agents center modeling, intervention, and evaluation on an evolving person and human agency rather than on external digital or physical task states. They combine longitudinal human-state understanding, prediction, proportionate intervention, and feedback under agency-preserving constraints.
- Core definition: Combodied Agents perceive, model, and influence a person’s evolving human state through continuous multimodal sensing and longitudinal interaction.Their perceptual and action domain includes the body, behavior, cognition, emotion, and surrounding context.
- Core definition: Their defining properties are human-centric state modeling, longitudinal reasoning, intervention, co-agency, and preservation of autonomy, control, dignity, relationships, and capability.These properties describe a center of gravity rather than a rigid interface category.
- Action substrates: Digital Agents primarily transform digital states, Embodied Agents transform physical or simulated physical states, and Combodied Agents support evolving human states while preserving agency over time.The categories overlap in implementation and may share perception, representation, reasoning, planning, action, learning, and evaluation machinery.
- Action substrates: A Combodied Agent evaluates whether support helps a person regulate health, understand decisions, sustain relationships, develop capabilities, or pursue valued goals over time.Software tools, sensors, robots, and communication can be action channels without being the agent’s ultimate target.
- Functional organization: The closed loop comprises human-state perception, longitudinal memory, personal world modeling, and intervention planning and delivery, with feedback updating subsequent decisions.Perception estimates current state; memory organizes temporal evidence; world modeling predicts trajectories under alternatives; intervention selects whether, when, and how to provide proportionate support or escalation.
- Functional organization: Human-centricity and longitudinality determine what the system models, while co-agency and agency preservation constrain the complete loop.The agent maintains uncertainty-bearing representations and retrieves decision-relevant evidence from personal memory, context, and goals.
3 Event-Based Multimodal Perception
Combodied Agents define multimodal perception as reconstructing meaningful, uncertain personal events from sparse, heterogeneous signals for longitudinal human-state understanding. The approach links acquisition context, temporal change, user reports, and modality-specific limitations while separating observations from events, inferred states, predictions, and intervention permission.
- Event-based perception: Event-based personal data perception filters, aligns, and interprets fragmentary data into evidence records for human-state understanding and longitudinal modeling.Records preserve what was observed, when and how it was acquired, the supported interpretation, remaining uncertainty, and relevance to memory or intervention.
- Acquisition context: Acquisition mode shapes the meaning, coverage, burden, privacy, and interpretation of evidence across user-reported, device-mediated, contact, ambient, and institutional configurations.The same nominal modality can differ substantially depending on whether it is actively authored, sensed on-body, captured ambiently, or obtained from an authorized record.
- Longitudinal evidence: Personal data perception targets transitions, deviations, social episodes, safety-relevant events, and intervention responses rather than isolated samples.Longitudinal evidence is sparse and discontinuous, so inferred states remain linked to observable fragments without becoming uncontestable facts.
- Modality limitations: Speech, vision, and physiology provide useful but indirect evidence whose interpretation depends on context, attribution, acquisition quality, and sensor reliability.Pauses, facial movement, missing data, and physiological readings can have multiple explanations; local processing, uncertainty tracking, and purpose-specific event descriptions help constrain interpretation.
- Inference boundaries: The architecture distinguishes observations from reconstructed events, inferred states, predicted trajectories, and permission to intervene.Text, physiological signals, speech, and vision are sources of event evidence rather than events themselves, and statistical confidence does not establish reliability, causal validity, or intervention permission.
4 Personal World Model
A Personal World Model (PWM) is a purpose-bounded, uncertainty-aware model that predicts how an individual’s state–event trajectory may unfold under alternative decisions and interventions. It informs scenario comparison while leaving authorization to an admissible intervention policy governed by consent, safety, reversibility, and agency preservation.
- Definition: A PWM predicts person-specific future states, events, and outcomes from governed multimodal history, current context, and candidate scenarios.Its outputs are uncertainty-bearing distributions rather than fixed forecasts.
- Definition: Its defining function is intervention-conditioned modeling of one person’s state–event trajectory across alternative scenarios, not personalization or plausible behavior generation alone.Subsequent events and intervention responses update the model’s representation of personal dynamics.
- Positioning: PWMs are families of purpose- and horizon-specific models whose components jointly support person-specific, context-sensitive, action-conditioned prediction with explicit uncertainty.Profiles, memory, personalized foundation models, simulations, causal models, and mechanistic models may implement parts of this contract.
- Predictive dynamics: Scenario rollouts vary unresolved user and environmental responses and compare distributions over future trajectories across non-intervention, clarification, and alternative support choices.Relevant outcomes include adherence, wellbeing, capability, safety, relationship quality, and agency preservation.
- Causal scope: A predictive PWM does not automatically identify individual counterfactuals: causal claims require explicit assumptions, defensible identification, and suitable intervention or observational designs.Writing do(·) alone does not remove confounding, and high-risk systems should not explore unconstrained actions merely to improve the model.
- Governance: The intervention policy, not the PWM, selects or authorizes actions within constraints on consent, scope, safety, uncertainty, reversibility, and escalation.High predicted benefit does not itself constitute permission to act.
5 From Cloud LLMs to Edge Personal Models
The paper proposes a three-stage deployment progression from cloud-centric systems to edge-mediated authority and finally user-controlled personal intelligence. The stages are architectural centers of gravity, evaluated by control, correctability, resilience, privacy, and human benefit rather than inference location alone.
- Stage I: Cloud-centric: Stage I places most reasoning, memory retrieval, tool selection, and response generation in the cloud, with the device serving mainly as interface, sensor endpoint, and execution surface.This architecture remains useful for prototyping, cold starts, broad knowledge, and low-risk applications.
- Stage I: Boundaries: Stage I reaches a governance boundary when service-controlled memory is treated as a persistent personal model or cloud outputs directly authorize irreversible actions.High-impact interventions require explicit confirmation or another trusted decision boundary, and corrections or deletions must propagate across affected representations.
- Stage II: Hybrid: Stage II makes the edge a privacy, interpretation, and authority mediator while retaining cloud support for external knowledge, complex reasoning, generation, and specialized tools.Raw signals and private memories can remain local while purpose-limited summaries are sent outward.
- Stage II: Boundaries: Hybrid systems must manage misclassified sensitivity, residual privacy leakage, inconsistent personal-state versions, and changing cloud interpretations of local summaries.These risks define the central technical boundary in dividing state and authority across components.
- Stage III: User control: Stage III keeps authoritative personal memory, the PWM, intervention policy, and safety boundaries primarily on trusted user-side devices while invoking cloud services selectively.Users should be able to inspect, correct, delete, pause, export, reset, and migrate represented model state.
- Comparison: The three stages should be compared by demonstrated control, correctability, resilience, and human benefit rather than nominal inference location.A single agent may occupy different stages for different tasks.
6 Benchmark & Evaluation
The paper argues that Combodied Agents require scenario-centered, longitudinal evaluation combining reliability, model quality, intervention appropriateness, human outcomes, and agency preservation. It proposes standardized benchmark episodes and the modular CombodiedBench suite, with severe safety and governance failures treated as non-compensatory.
- Evaluation scope: Task completion and physical safety are necessary but insufficient when agents can alter a person’s state, capability, relationships, or safety over time.Evaluation must jointly test system reliability, model quality, intervention appropriateness, human outcomes, and agency preservation.
- Scenario-centered evaluation: Evaluation should begin from the scenario and authority level because the same intervention can help adherence, undermine autonomy, or deepen dependency in different contexts.The matrix must therefore be instantiated differently for each scenario.
- Agency preservation: Agency preservation measures whether repeated support protects and strengthens the user’s capacity to understand, choose, act, refuse, correct, and grow.It is distinct from task success, satisfaction, personalization quality, and engagement.
- Agency metrics: Agency evaluation should report multiple dimensions across interaction, episode, and longitudinal levels rather than a single scalar score.Metrics include autonomy preservation, contestability, correction, and scenario-appropriate authority standards.
- Agency metrics: Baselines should compare users with and without agent support or compare direct execution, confirmation-based execution, reflective coaching, and no intervention.The key question is whether assistance changes future ability, not only immediate success.
- Benchmark construction: A benchmark episode should include prior trajectory, current context, available evidence, permissible actions, delayed outcomes, authority, explanations, horizon, scoring, and unacceptable failures.Scoring combines automatic checks, expert judgment, and longitudinal outcomes, while critical failures remain non-compensatory.
- CombodiedBench: CombodiedBench is a modular suite covering state perception, memory continuity, goal negotiation, intervention appropriateness, agency, relationship boundaries, escalation, and longitudinal outcomes.Its multidimensional results should document scenario, label, safety-case, and evaluator changes to preserve comparability.
7 Taxonomy and Applications
Combodied Agents are organized by the human state they target, the relationship in which the user is situated, and the role the agent adopts. This taxonomy is longitudinal and links each category to distinct permissions, interventions, safeguards, and agency outcomes.
- Three-axis taxonomy: The taxonomy uses three axes: human-state target, relational context, and agent role.These axes classify what the agent changes, the social position from which it acts, and the role it takes within that relationship.
- Human-state targets: Human-state targets determine the signals, memories, intervention policies, evaluation criteria, and safety boundaries a system requires.Targets include cognitive support, habit change, health care, emotional support, life management, protection, and identity reflection.
- Agent roles: Agent roles are contextual and may shift among tool, coach, mediator, caregiver, and advocate, but transitions must be explicit and visible.Role clarity prevents a scheduling tool from silently becoming a behavioral coach or a caregiver from becoming a surveillance system.
- Longitudinal evaluation: The taxonomy is longitudinal: evaluation must track state dynamics, accumulated dependency or harm, progress, and when oversight is required.Cross-cutting dimensions include memory scope, initiative, intervention intensity, evaluation targets, and deployment locus.
- Relational contexts: Relationship mode changes the norms, obligations, vulnerabilities, permissions, memory scope, intervention rights, and represented interests governing support.A single global memory is unsafe when family, health, intimate, or workplace contexts require scoped memories.
- Design implication: Combodied Agents are defined by what aspect of a person’s life-world they model, how they act within relationships, and whether interventions preserve or strengthen agency over time.Sensitive domains often require local-first architectures and stricter cloud-routing controls than generic productivity support.
8 Risks, Challenges, and Future Directions
Combodied Agents create risks because persistent personalization can influence people, access sensitive life data, and shape agency over time. The research agenda therefore emphasizes bounded control, safer infrastructure, personal-dynamics modeling, and long-term accountability.
- Agency risks: Persistent personalized agents can produce manipulation, dependency, sycophancy, and business-model misalignment when immediate compliance or engagement displaces long-term agency.These risks arise from the agent’s ability to understand and influence users across time.
- Agency risks: Manipulation risk increases when agents exploit moments of loneliness, fatigue, anxiety, or cognitive overload.Governance should constrain commercially motivated nudging and require transparent purposes, user control, and auditability for high-impact actions.
- Privacy and control: Privacy governance must control what agents sense, remember, infer, share, recommend, and execute.Users need correction, forgetting, minimized retention, distinctions between facts and inferences, and audit logs for sensitive access.
- Privacy and control: Consent should define permissions and boundaries, while override must allow users to refuse, pause, correct, delete, revoke, or reverse agent actions and memories.High-impact actions require confirmation, high-risk states require escalation, and initiative must remain bounded by meaningful human control.
- High-risk settings: Vulnerable users and medical or mental-health settings require stricter defaults, professional boundaries, uncertainty communication, crisis protocols, and reliable escalation.In high-risk situations, safe behavior may require handoff to qualified human support rather than better conversation.
- Future directions: Open problems include learning personal dynamics and intervention effects from sparse longitudinal evidence, translating uncertainty into agency-aligned support, and building trusted personal infrastructure.The agenda treats progress as increased capability that remains corrigible, culturally situated, accountable, and aligned with long-term agency.
9 Conclusion
The paper presents Combodied Agents as a human-centric Agentic AI paradigm centered on supporting human trajectories without undermining autonomy, capability, safety, or relationships. It proposes judging progress by whether people retain and develop the ability to understand, choose, correct, recover, and sustain meaningful lives.
- Combodied Agents unify a closed-loop framework, taxonomy, deployment perspective, and evaluation agenda for human-centered Agentic AI.
- Their central design principle is to act with users in ways that preserve and strengthen long-term agency.
- Progress should be judged by whether people can understand, choose, correct, recover, develop capability, sustain relationships, and live according to evolving values.