Source-linked AI summary
Human-Centric Intelligence in the Era of Foundation Models: A Survey
Yang Chen, Tianqi Wang, Xiaorui Jiang, Yilei Man, Yihua Shao, Mengyuan Liu, Zhi Chen, Xiaofeng Cao, Qibin Zhao, Chi Harold Liu, Albert Y. Zomaya, Nicu Sebe, Jingren Zhou, Dacheng Tao, Song Guo, Jingcai Guo
TL;DR
Human-centric intelligence remains fragmented across tasks, modalities, and research communities and has not fully integrated with foundation models. This survey addresses the gap with a six-level human-context taxonomy and systematic review, concluding that the field is moving toward scalable data, reusable priors, multimodal interfaces, and transferable capabilities.
Problem
Human-centric intelligence remains difficult to organize because advances span fragmented tasks, modalities, and research communities without clear conceptual and methodological connections.
Method
The survey introduces a six-level human-context taxonomy and organizes methods, data families, architectures, training strategies, datasets, benchmarks, and evaluation metrics.
Results
The analysis characterizes the foundation-model era as a transition toward scalable human data, reusable priors, multimodal interfaces, and transferable capabilities.
Takeaways & Limitations
The taxonomy and organized evidence provide a coherent framework and practical reference for advancing human-centric intelligence across its six interconnected levels.
Takeaways & Limitations
Perceptual accuracy in foundation world models does not ensure decision utility, which also requires reliable interventions, consequential-change separation, and calibrated uncertainty.
Abstract
from arXiv · showhide
Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, and general-purpose modeling. Yet it has not fully integrated with foundation models to achieve the comparable progress seen in them. More importantly, recent advances across this broad landscape remain fragmented across tasks, modalities, and research communities, leaving their intrinsic conceptual and methodological connections unclear. To bridge these divides and rethink human-centric intelligence in the foundation-model era, we introduce a full-spectrum human context taxonomy that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency. We next present the methodological foundations of the field, covering human-centric data families, computational architecture paradigms, and representative training and inference optimization strategies. We then systematically review representative methods across these levels and organize the associated datasets, benchmarks, and evaluation metrics. We further discuss open challenges and promising research directions toward human-centric intelligence that is scalable, trustworthy, physically grounded, and deployable, aiming to provide a coherent framework and practical reference for advancing the field. Finally, we provide a systematically organized and continuously updated collection of human-centric AI literature and resources on our project page.
1 Introduction
This survey reframes human-centric intelligence in the foundation-model era as a full-spectrum field spanning observable subjects, dynamic actors, and situated agents. It unifies six human-context levels with methodological foundations, evaluation protocols, open challenges, and organized resources.
- The Foundation-Model Era Shift: It addresses a foundation-model-era shift from hand-crafted descriptors and dataset-specific deep learning [272] toward broader scalable and transferable human modeling.The survey includes methods using scalable learning, reusable pretraining, multimodal interfaces, cross-task generalization, or transferable human priors.
- Comparison with Other Human-Centric Surveys: Unlike task-focused surveys of pose estimation, action recognition, motion and video generation [72], and human-object interaction [212], this survey connects human contexts across tasks.Its broader organization responds to fragmentation among human-centric research areas and emphasizes their relationships.
- Scope and Contributions: The survey organizes human-centric intelligence into six connected levels: visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, and embodied agency.The taxonomy progressively expands from visual observation and body structure to motion, interactions, world dynamics, and physical execution.
- Scope and Contributions: The survey supplies methodological foundations covering human-centric data families, computational architecture paradigms, and training and inference optimization strategies.These foundations guide its review of representative methods across the human-context taxonomy.
- Scope and Contributions: It organizes representative datasets, benchmarks, and evaluation metrics to clarify how different human-centric capabilities are empirically developed and assessed.These evaluation protocols provide a consolidated empirical basis for comparing research across the field.
- Scope and Contributions: The survey discusses open challenges and future directions and releases systematically organized resources to support continued progress in human-centric AI.The project page provides a continuously updated collection of human-centric AI literature and resources.
2 Human Context Taxonomy
The taxonomy organizes fragmented human-centric intelligence into six interconnected levels across three perspectives: observable subjects, dynamic actors, and situated agents. It assigns multi-context methods by their primary modeling target and evaluation objective, spanning appearance and geometry, dynamics and interaction, and world simulation and embodied agency.
- Taxonomy Overview: The six interconnected levels are visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, and embodied agency, organized across three perspectives on humans.Methods spanning multiple contexts are assigned to the primary level defined by their modeling target and evaluation objective.
- Observable Subjects: Observable-subject modeling separates visual appearance, focused on visible and identity-bearing properties, from spatial geometry, which represents pose, shape, and body organization across viewpoints.Together, these levels provide perceptual and structural foundations for broader human-centric capabilities and connect visual perception with physical human modeling.
- Observable Subjects: Visual-appearance tasks include generalist perception, identity recognition and retrieval, and controllable human generation, while spatial-geometry tasks include pose and mesh recovery and renderable avatar modeling.Examples range from pose estimation and person re-identification to identity-preserving generation, parametric body recovery, animatable avatars, and novel-view rendering.
- Dynamic Actors: Dynamic-actor modeling distinguishes kinematic dynamics, which captures intrinsic temporal body evolution, from interaction modeling, which incorporates objects, scenes, and other people as behavioral counterparts.Kinematic tasks cover motion understanding, generation, and human video animation; interaction tasks cover human-object, human-scene, and social interaction.
- Situated Agents: Situated-agent modeling separates world simulation, which co-evolves human actions and world states, from embodied agency, which turns human-centered knowledge into executable behavior.These levels extend the taxonomy from describing humans and their relations toward predicting environmental consequences and enabling physical or operational action.
3 Preliminary
Human-centric intelligence in the foundation-model era rests on heterogeneous human-centered data, computational architectures, and optimization strategies that together shape learned information, capabilities, and their acquisition, adaptation, and elicitation.
- Heterogeneous human-centered data define what information human-centric systems can learn.
- Computational architectures determine how human-centered information is transformed into capabilities.
- Optimization strategies govern how capabilities are acquired, adapted, and elicited across specialized and general-purpose pretrained models.These methodological foundations apply to human-specialized foundation models and methods adapting general-purpose pretrained models.
3.1 Human-Centric Data and Signal Families
The section organizes human-centric data into five signal families based on acquisition mechanisms and preserved information, spanning visual, geometric, sensorimotor, wireless/ranging, and linguistic/acoustic observations. These families offer complementary strengths and limitations for scalable, multimodal human-centric intelligence.
- Taxonomy: Five signal families organize human-centric data by acquisition mechanism and preserved information: visual, spatial-structural, sensorimotor, wireless/ranging, and linguistic/acoustic signals.This taxonomy is illustrated in Fig. 4 and structures the principal signals and representation formats used throughout the field.
- Visual imaging: Visual imaging provides semantically rich and scalable observations through exocentric and egocentric RGB, infrared and thermal, and event-camera signals, [93],,.These signals support broad capability development but remain vulnerable to occlusion, viewpoint and illumination changes, motion blur, and privacy concerns.
- Sensorimotor signals: Sensorimotor signals cover gaze, inertial measurements [4], action trajectories [393], control signals [297], tactile contact [66], and physiological measurements [357].They provide temporally precise and potentially privacy-preserving information unavailable from vision alone, while scaling is limited by placement, calibration, drift, variation, synchronization, and data scarcity.
- Wireless and ranging signals: Wireless and ranging signals, including WiFi [4] [69], radar, and LiDAR, provide non-contact sensing when RGB imaging is unreliable or undesirable.They offer limited semantic detail and remain sensitive to sensor configuration, environmental conditions, and transfer across sensing systems.
- Linguistic and acoustic signals: Text and audio provide semantic and communicative interfaces for cross-modal understanding, controllable generation, and instruction-conditioned modeling.They describe behavior beyond directly observable body states but indirectly capture geometry and motion, with informativeness depending on semantic specificity, recording quality, transcription accuracy, and temporal alignment.
3.2 Computational Architectures
Section 3.2 organizes human-centric foundation-model architectures into four paradigms according to how encoded inputs are transformed into capability-specific outputs. These paradigms range from single-step mappings to sequential, iterative, and hybrid computations.
- Single-step mapping architectures: Single-step mapping architectures transform inputs into target representations or structured outputs through one feedforward computation, including encoder-only, encoder-decoder, and JEPA-style designs.Encoder-only models produce human representations, encoder-decoder models generate structured predictions, and JEPA-style models estimate target embeddings.
- Sequential factorization architectures: Sequential factorization architectures model outputs as ordered sequences using conditional next-token prediction, supporting text, motion, and actions [297] while capturing long-range dependencies.Autoregressive LLM and MLLM architectures apply this formulation to linguistic symbols or discretized human representations.
- Iterative generation architectures: Iterative generation architectures progressively transform an initial state toward a target distribution, with diffusion using reverse denoising and flow-based models [346] learning continuous transport fields.Unlike token-by-token generation, both operate on the output as a globally evolving state.
- Hybrid architectures: Hybrid architectures combine at least two computational mechanisms through sequential, shared-representation, parallel, or iterative-feedback compositions to produce the final output.Their mechanisms may be integrated through a composition function that combines intermediate representations.
3.3 Optimization Strategies
Optimization strategies govern how human-centric capabilities are acquired during training and elicited during inference. Training changes all or selected parameters, whereas inference keeps learned parameters fixed and improves test-time computation, with both stages supporting combinable mechanisms.
- 3.3.2 Inference: Inference optimization keeps model parameters fixed and improves capability elicitation through semantic augmentation, guided sampling, iterative refinement, preference-based selection, and retrieval augmentation.These mechanisms may be combined within one inference process.
- 3.3.1 Training: Training includes seven mechanisms—scratch-based training, full-parameter fine-tuning, parameter-efficient tuning, instruction tuning, reward-based tuning, knowledge distillation, and task-head tuning—that can be combined across multi-stage pipelines.These mechanisms differ in initialization, updated parameters, and optimization signals.
- 3.3.1 Training: Training strategies trade off capability acquisition, adaptation, preservation of general knowledge, computational cost, and transferability under different data and resource constraints.Scratch training requires large-scale data and computation; full fine-tuning supports substantial domain adaptation but may weaken general capabilities, while parameter-efficient and task-head tuning reduce adaptation cost and can test feature transfer.
- 3.3.1 Training: Instruction, reward-based, and distillation training improve controllability, behavior alignment, and capability transfer when direct supervised targets are insufficient or capabilities span architectures, modalities, or pipeline stages.Instruction tuning exposes capabilities through a shared language-facing interface; reward-based tuning uses evaluated outputs or reinforcement-learning-style objectives; distillation transfers outputs, states, or behaviors while reducing inference cost.
- 3.3.2 Inference: Inference mechanisms enrich inputs, steer generation, correct outputs, select among candidates, or retrieve external knowledge to improve ambiguity, controllability, consistency, physical validity, quality, and instance-level human-centric reasoning.Iterative refinement addresses geometry, temporal coherence, interaction feasibility, and physical validity; retrieval augmentation adds relevant examples, contexts, or priors without changing parameters.
4 Human Visual Appearance and Spatial Geometry
This section frames visual appearance and spatial geometry as the observable-subject perspective of human-centric intelligence: appearance captures image-space cues, while geometry represents the body as explicit or renderable structure. In the foundation-model era, both are moving toward reusable human priors and flexible interfaces, with appearance organized around generalist perception, identity understanding, and controllable generation, and geometry around structured modeling and renderable avatars.
- Generalist human perception: Human-centric appearance methods are shifting from task-specific models toward specialized pretraining and unified interfaces that transfer reusable knowledge across perceptual tasks.HumanBench, HAP, Sapiens, DAViD, THFM [324], Hulk, and facial models illustrate this movement; expanding task coverage without sacrificing performance remains a central challenge.
- Discriminative identity understanding: Identity understanding is moving from dataset-specific matching toward recognition across heterogeneous observations and queries, but stable identity requires separating persistent traits from changing appearance cues.Recent models address face-and-body variation, instruction-guided retrieval, multimodal queries, structured identity supervision, and language-based grounding [108] [109] [64] [140]. Privacy, demographic variation, uncertainty, and adaptability remain important constraints.
- Controllable human generation: Controllable human generation is advancing from generic synthesis toward compositional control over identity, anatomy, garments, pose, editing, and animation [44] [52] [12].The key difficulty is editing targeted human properties while preserving unrelated characteristics, requiring reusable foundation architectures that combine multiple intentions without interference.
- Spatial geometry: Spatial geometry extends human-centric intelligence beyond image-space appearance by modeling explicit body structure or renderable avatar assets through scalable human priors and transferable semantic interfaces.The subsection distinguishes structured geometry modeling from renderable avatar modeling, which integrates geometry and appearance into renderable assets.
5 Human Kinematic Dynamics and Interaction Modeling
This section presents kinematic dynamics and interaction modeling as the dynamic-actor perspective of human-centric intelligence, covering temporal body evolution and relations with objects, environments, and people. It organizes kinematic dynamics into scalable motion modeling and human video animation, while framing interaction modeling across human-object, human-scene, and social interaction.
- Section overview: Kinematic dynamics and interaction modeling frame humans as dynamic actors whose temporal body evolution is shaped by relations with objects, environments, and other people.The field is moving beyond isolated motion generators and interaction-specific pipelines toward reusable temporal priors, multimodal interfaces, and more general models.
- Human video animation: Human video animation realizes body dynamics in photorealistic video through audio-, pose-, motion-, appearance-, and multimodal conditioning, increasingly unifying video generation with structured motion.Large video backbones improve realism and flexibility, but precise driving-to-generated-body correspondence, long-horizon coherence, identity consistency, anatomical plausibility, and subtle timing remain unresolved.
- Scalable motion modeling: Scalable motion modeling treats movement as a temporal signal, advancing reusable temporal and linguistic interfaces, motion-language unification, multimodal observation, and tool- or retrieval-assisted control.Representative systems include MotionBERT for transferable masked-sequence features, MotionGPT for discrete-token captioning and generation [133], AvatarGPT for planning and task decomposition, and MG-MotionLLM for shared motion understanding and generation.
- Scalable motion modeling: Motion-native foundation models increasingly study scaling and transfer across data, architectures, and tasks, but current scaling captures physical validity, temporal causality, and subtle human variation only partially.Examples include Being-M0 [332], ScaMo, Go-to-Zero [70], Kimodo [283], and HY-Motion [346]; the central challenge is preserving continuous kinematics across objectives and observation domains.
- Interaction modeling: Interaction modeling expands from local human-object relations to human-scene compatibility and social coordination through open-ended semantic knowledge, multimodal reasoning, and reusable generative priors.This organization treats humans as relational actors whose behavior is jointly shaped by external entities.
6 Human World Simulation and Embodied Agency
The section frames world simulation and embodied agency as the situated-agent levels of human-centric intelligence, distinguishing perceptual world generation from actionable planning and executable physical behavior. It reviews foundation-model approaches while emphasizing that useful systems require persistent, causally responsive, controllable worlds and physically grounded agency.
- World Simulation: World simulation separates human-centered world generation, which models perceptual futures around human activity, from actionable world planning, which predicts action-dependent transitions for decisions and embodied action.This distinction is illustrated in Fig. 10 and separates visually plausible simulation from predictive models whose outputs or internal variables support action.
- World Simulation: Foundation video models adapt human motion, viewpoint, goals, scene structure, and full-body trajectories into increasingly persistent first-person simulations, but their outputs remain perceptual rather than executable.Examples include Generated Reality [361], Hand2World [339], AnchorWorld [199], EgoForge [293], EgoControl, EgoX, and EgoWorld [266].
- World Simulation: Actionable world models condition future prediction on pose, manipulation, or dexterous actions and increasingly couple visual forecasting with action generation, interaction learning, and closed-loop planning.Representative systems include PEVA [8], LOME [87], EgoHOI [175], HandWorld, DexWM [91], EgoAgent [43], EgoSim [106], and EgoExo-WM [315].
- World Simulation: Perceptual accuracy alone does not guarantee decision utility: actionable world models require reliable intervention responses, separation of consequential changes from incidental variation, and calibrated uncertainty.The relationship between generative fidelity and planning effectiveness remains a central challenge in human-centered world modeling.
- Embodied Agency: Embodied agency distinguishes generalist humanoid control from human-to-agent skill transfer, using foundation models as semantic engines alongside reusable motor priors, specialized controllers, and scalable behavioral pretraining.Representative directions include VGHuman [392], HOI-HLI, BiBo [132], VLM-RMD [60], MaskedMimic, InterMimic, TokenHSI, SCRIPT [414], SONIC [240], and Humanoid-GPT [275].
7 Datasets, Benchmarks, and Metrics
This section reviews representative datasets, benchmarks, and commonly used metrics that support human-centric model development and evaluation, while identifying limitations in current evaluation practices.
- 7 Datasets, Benchmarks, and Metrics: Datasets and benchmarks define what human-centric models learn from and how their capabilities are evaluated.The section reviews representative resources before organizing commonly used metrics.
- 7 Datasets, Benchmarks, and Metrics: Organized metrics clarify current evaluation practices and reveal where those practices remain incomplete.
7.1 Datasets and Benchmarks
Datasets provide observations and annotations, while benchmarks standardize tasks and protocols for comparing human-centric model capabilities. The resources are organized into four taxonomy-aligned groups—human subjects, human dynamics, human interactions, and human embodiment—covering progressively richer perceptual, temporal, relational, and embodied contexts.
- Organization: The resource taxonomy groups datasets and benchmarks into human subjects, human dynamics, human interactions, and human embodiment, linking heterogeneous resources to the progression of human context.Datasets support model development through observations and annotations, whereas benchmarks enable standardized comparison through defined tasks and protocols.
- Human Subjects: Human subject resources progress from visual and multisensory observation to identity correspondence and renderable 3D or 4D human geometry.They support appearance modeling, multimodal sensing under occlusion or privacy constraints, person-level matching, reconstruction, animation, relighting, and controllable synthesis.
- Human Dynamics: Human dynamics resources cover video generation and animation, behavior understanding, kinematic motion, gait, sport analysis, and virtual try-on through temporally structured or conditioned data.They increasingly support fine-grained motion reasoning, motion-language supervision, recognition across changing observation conditions, athletic feedback, and identity-preserving garment or footwear transfer.
- Human Interactions: Human interaction resources represent behavior in procedural, physical, environmental, and social contexts, organized by egocentric activities, human-object interaction, human-scene interaction, and social interaction.Egocentric procedural datasets support step recognition, intent anticipation, error detection, and assistance from the actor’s viewpoint, with increasingly rich hand, assistance, and dense 3D annotations.
7.2 Metrics
The section organizes human-centric evaluation into five metric families—reconstruction-centric, semantic-centric, generative-centric, interaction-centric, and efficiency-centric—because heterogeneous outputs span geometry, semantics, generated content, physical interactions, and executable behavior. It further structures these metrics by increasingly broad evaluation targets, from output agreement and semantic preservation to generative quality and contextual validity.
- Metric Taxonomy: Human-centric evaluation is classified into reconstruction-centric, semantic-centric, generative-centric, interaction-centric, and efficiency-centric metrics to cover heterogeneous outputs and evaluation targets.The taxonomy spans geometric structures, semantic predictions, generated content, physical interactions, and executable behavior.
- Reconstruction-Centric Metrics: Reconstruction-centric metrics measure agreement with reference observations across articulated human state, surface geometry, and image or rendering fidelity.Examples include MPJPE and PA-MPJPE for joint accuracy, PVE/MPVPE for dense meshes, Chamfer Distance and Surface F-score for surfaces, and PSNR, SSIM, and LPIPS for visual fidelity.
- Semantic-Centric Metrics: Semantic-centric metrics evaluate task-relevant, identity-related, and cross-modal meaning through recognition, retrieval and verification, and semantic alignment.They include accuracy, precision, recall, F1, AP/mAP, Rank-k, CMC, mINP, R-Precision, Recall@K, CLIP- and DINO-based similarities, and language-generation metrics.
- Generative-Centric Metrics: Generative outputs require complementary metrics for fidelity, diversity and coverage, temporal dynamics, audio-visual synchronization, and identity or appearance consistency.These dimensions use distributional measures such as FID and FVD, diversity and coverage measures, trajectory and smoothness diagnostics, synchronization scores, and identity or garment consistency measures.
- Interaction-Centric Metrics: Interaction-centric metrics assess whether predicted or generated behavior remains valid across contact and collision, physical plausibility, embodied task performance, and human preference or subjective quality.The organization progresses from local geometric validity to physical execution, task completion, and perceived naturalness.
8 Open Challenges and Future Directions
The field is moving toward general human-centric systems, but scaling individual tasks alone cannot resolve challenges in data, unified representation, physical grounding, world modeling, evaluation, and deployment. Six directions target scalable, transferable, trustworthy, physically grounded, and efficient human-centric intelligence.
- Data Scaling: Limited real human data motivates synthetic-data scaling, enabled by automatic annotation, low marginal cost, and targeted coverage of difficult real-world cases.Data collection requires specialized capture systems, extensive annotation, and careful treatment of personal information.
- Data Scaling: Future studies should separate data quantity from quality, test real-synthetic mixtures on real and underrepresented populations, and preserve provenance and consent through adaptation [9].These protocols could clarify when synthetic data provides genuine transfer.
- Unified Foundation Models: A general human-centric foundation model should replace modality- or task-restricted systems with unified representation learning across the full human context spectrum [12] [324].Current unified models typically cover restricted task groups or particular modalities.
- Physical Grounding: Next-generation representations should model physical information centrally, because visual and linguistic regularities can support plausible behavior without capturing the physical laws enabling it [35] [92] [122].Physical grounding should not be applied only as a post-prediction or post-generation constraint.
- World Models: Human-centric world models should jointly predict agent actions and environmental responses, since visually plausible futures do not establish causal action-consequence modeling [293] [339] [361].Human videos offer scalable experience, but the embodiment gap limits direct transfer from observed behavior.
- Evaluation: Evaluation should connect benchmarks across perception, generation, and interaction, because separate task performance does not establish cross-context transfer or capability composition [28] [323].The central limitation is the lack of protocols linking existing metrics, not the absence of metrics themselves.
- Deployment: Efficient deployment requires alternatives to monolithic models that address training cost, inference latency, privacy exposure, and maintenance complexity through agentic, tool-augmented coordination.Such systems can coordinate foundation models, specialized human models, and external perception tools.
9 Conclusion
This survey presents a full-spectrum account of human-centric intelligence in the foundation-model era, organized by a six-level human context taxonomy and methodological foundations.
- Conclusion: The survey presents a full-spectrum account of human-centric intelligence in the foundation-model era.
- Conclusion: Its human context taxonomy connects visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, and embodied agency.
- Conclusion: The taxonomy frames humans as observable subjects, dynamic actors, and situated agents while guiding the field’s methodological foundations and review of advances.