Source-linked AI summary
Human-Centric Intelligence in the Era of Foundation Models: A Survey
Yang Chen, Tianqi Wang, Xiaorui Jiang, Yilei Man, Yihua Shao, Mengyuan Liu, Zhi Chen, Xiaofeng Cao, Qibin Zhao, Chi Harold Liu, Albert Y. Zomaya, Nicu Sebe, Jingren Zhou, Dacheng Tao, Song Guo, Jingcai Guo
TL;DR
Human-centric intelligence는 과제, modality, 연구 공동체 전반에 걸쳐 여전히 파편화되어 있으며 foundation models와도 충분히 통합되지 않았다. 이 survey는 six-level human-context taxonomy와 체계적 문헌 검토를 통해 이러한 간극을 다루고, 분야가 scalable human data, 재사용 가능한 priors, multimodal interfaces, transferable capabilities를 향해 나아가고 있다고 결론짓는다.
문제
Human-centric intelligence는 명확한 개념적·방법론적 연결 없이 발전이 파편화된 과제, modality, 연구 공동체에 걸쳐 이루어지기 때문에 체계화하기 여전히 어렵다.
방법
이 survey는 six-level human-context taxonomy를 제시하고, 방법론, data family, architecture, training strategy, dataset, benchmark, evaluation metric을 체계적으로 구성한다.
결과
분석에 따르면 foundation-model era는 scalable human data, 재사용 가능한 priors, multimodal interfaces, transferable capabilities를 향한 전환기다.
시사점 및 한계
이 taxonomy와 체계적으로 정리된 근거는 상호 연결된 six levels 전반에서 human-centric intelligence를 발전시키기 위한 일관된 framework이자 실용적 reference를 제공한다.
시사점 및 한계
foundation world models의 지각 정확도가 decision utility를 보장하지는 않으며, 이를 위해서는 신뢰할 수 있는 intervention, consequential change의 분리, calibrated uncertainty도 필요하다.
Abstract
from arXiv · showhide
Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, and general-purpose modeling. Yet it has not fully integrated with foundation models to achieve the comparable progress seen in them. More importantly, recent advances across this broad landscape remain fragmented across tasks, modalities, and research communities, leaving their intrinsic conceptual and methodological connections unclear. To bridge these divides and rethink human-centric intelligence in the foundation-model era, we introduce a full-spectrum human context taxonomy that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency. We next present the methodological foundations of the field, covering human-centric data families, computational architecture paradigms, and representative training and inference optimization strategies. We then systematically review representative methods across these levels and organize the associated datasets, benchmarks, and evaluation metrics. We further discuss open challenges and promising research directions toward human-centric intelligence that is scalable, trustworthy, physically grounded, and deployable, aiming to provide a coherent framework and practical reference for advancing the field. Finally, we provide a systematically organized and continuously updated collection of human-centric AI literature and resources on our project page.
1 서론
이 survey는 foundation-model 시대의 human-centric intelligence를 관찰 가능한 대상, 동적 행위자, 상황 속 agent를 아우르는 full-spectrum 분야로 재구성한다. 또한 여섯 human-context 수준을 방법론적 기반, 평가 프로토콜, 미해결 과제, 체계화된 resource와 통합한다.
- Foundation-Model 시대의 변화: foundation-model 시대의 변화에 따라 hand-crafted descriptor 와 dataset-specific deep learning [272]에서 더 넓은 범위의 확장 가능하고 전이 가능한 human modeling으로 나아가는 흐름을 다룬다.확장 가능한 learning, 재사용 가능한 pretraining, multimodal interface, cross-task generalization, transferable human prior를 활용하는 방법을 포함한다.
- 다른 Human-Centric Survey와의 비교: pose estimation, action recognition, motion and video generation [72], human-object interaction [212]에 초점을 둔 survey와 달리, 과제 전반의 human context를 연결한다.더 넓은 구성은 human-centric 연구 영역 간의 fragmentation에 대응하며 이들의 관계를 강조한다.
- 범위와 기여: human-centric intelligence를 visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, embodied agency라는 서로 연결된 여섯 수준으로 구성한다.이 taxonomy는 visual observation과 신체 구조에서 motion, interaction, world dynamics, physical execution으로 점진적으로 확장된다.
- 범위와 기여: human-centric data family, computational architecture paradigm, training 및 inference optimization strategy를 아우르는 방법론적 기반을 제공한다.이러한 기반은 human-context taxonomy 전반의 대표 방법을 검토하는 지침이 된다.
- 범위와 기여: 대표 dataset, benchmark, evaluation metric을 체계화하여 서로 다른 human-centric capability가 경험적으로 어떻게 개발되고 평가되는지 명확히 한다.이러한 evaluation protocol은 분야 전반의 연구를 비교하기 위한 통합된 경험적 기반을 제공한다.
- 범위와 기여: 미해결 과제와 미래 방향을 논의하고, human-centric AI의 지속적인 발전을 지원하기 위해 체계적으로 정리된 resource를 공개한다.project page는 human-centric AI 문헌과 resource를 지속적으로 갱신하여 제공한다.
2 인간 맥락 Taxonomy
이 taxonomy는 인간 중심 지능의 파편화된 연구를 관찰 가능한 주체, 동적 행위자, 상황화된 에이전트라는 세 관점에 걸친 여섯 개의 상호 연결된 수준으로 구성한다. 다중 맥락 방법은 주요 modeling target과 evaluation objective에 따라 분류하며, appearance와 geometry, dynamics와 interaction, world simulation과 embodied agency를 포괄한다.
- Taxonomy 개요: 여섯 개의 상호 연결된 수준은 visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, embodied agency이며, 인간에 대한 세 관점에 걸쳐 구성된다.여러 맥락을 포괄하는 방법은 modeling target과 evaluation objective에 의해 정의되는 주요 수준에 배정된다.
- 관찰 가능한 주체: 관찰 가능한 주체에 대한 modeling은 가시적이며 정체성을 담는 속성에 초점을 둔 visual appearance와, 여러 viewpoint에서 pose, shape, body organization을 표현하는 spatial geometry를 구분한다.이 수준들은 함께 더 폭넓은 인간 중심 역량을 위한 지각적·구조적 기반을 제공하며, visual perception과 physical human modeling을 연결한다.
- 관찰 가능한 주체: Visual-appearance task에는 generalist perception, identity recognition and retrieval, controllable human generation이 포함되며, spatial-geometry task에는 pose and mesh recovery와 renderable avatar modeling이 포함된다.예로는 pose estimation과 person re-identification, identity-preserving generation, parametric body recovery, animatable avatars, novel-view rendering 등이 있다.
- 동적 행위자: 동적 행위자에 대한 modeling은 고유한 시간적 신체 변화를 포착하는 kinematic dynamics와, 객체·장면·다른 사람을 행동의 상대방으로 포함하는 interaction modeling을 구분한다.Kinematic task는 motion understanding, generation, human video animation을 다루며, interaction task는 human-object, human-scene, social interaction을 다룬다.
- 상황화된 에이전트: 상황화된 에이전트에 대한 modeling은 인간의 행동과 world state를 함께 변화시키는 world simulation과, 인간 중심 지식을 실행 가능한 행동으로 전환하는 embodied agency를 구분한다.이 수준들은 taxonomy를 인간과 인간의 관계를 기술하는 데서 환경적 결과를 예측하고 물리적 또는 운영적 행동을 가능하게 하는 방향으로 확장한다.
3 사전 지식
파운데이션 모델 시대의 인간 중심 지능은 서로 함께 학습되는 정보와 역량, 그리고 그 습득·적응·이끌어내기를 형성하는 이질적인 인간 중심 데이터, 계산 아키텍처, 최적화 전략에 기반한다.
- 이질적인 인간 중심 데이터는 인간 중심 시스템이 학습할 수 있는 정보를 규정한다.
- 계산 아키텍처는 인간 중심 정보가 역량으로 변환되는 방식을 결정한다.
- 최적화 전략은 전문화된 사전학습 모델과 범용 사전학습 모델 전반에서 역량이 습득·적응·이끌어내지는 방식을 좌우한다.이러한 방법론적 토대는 인간 특화 파운데이션 모델과 범용 사전학습 모델을 적응시키는 방법에 적용된다.
3.1 Human-Centric 데이터 및 신호 패밀리
이 절은 획득 메커니즘과 보존된 정보를 기준으로 human-centric 데이터를 시각, 기하, sensorimotor, wireless/ranging, 언어/음향 관측을 아우르는 다섯 신호 패밀리로 구성한다. 이 패밀리들은 확장 가능한 multimodal human-centric intelligence를 위해 상호보완적인 강점과 한계를 제공한다.
- Taxonomy: 다섯 신호 패밀리는 획득 메커니즘과 보존된 정보에 따라 human-centric 데이터를 구성한다: visual, spatial-structural, sensorimotor, wireless/ranging, linguistic/acoustic 신호다.이 taxonomy는 Fig. 4에 제시되며, 분야 전반에서 사용되는 주요 신호와 representation format을 체계화한다.
- Visual imaging: Visual imaging은 exocentric 및 egocentric RGB, infrared 및 thermal, event-camera 신호를 통해 의미적으로 풍부하고 확장 가능한 관측을 제공한다, [93],,.이러한 신호는 폭넓은 capability 개발을 지원하지만 occlusion, viewpoint 및 illumination 변화, motion blur, privacy 문제에 취약하다.
- Sensorimotor signals: Sensorimotor 신호는 gaze, inertial measurement [4], action trajectory [393], control signal [297], tactile contact [66], physiological measurement [357]를 포함한다.이 신호들은 vision만으로는 얻을 수 없는 시간적으로 정밀하고 잠재적으로 privacy-preserving한 정보를 제공하지만, 배치, calibration, drift, variation, synchronization, data scarcity로 인해 확장성이 제한된다.
- Wireless and ranging signals: WiFi [4] [69], radar, LiDAR 를 포함하는 Wireless 및 ranging 신호는 RGB imaging이 신뢰하기 어렵거나 바람직하지 않을 때 비접촉 sensing을 제공한다.이 신호들은 제한적인 semantic detail을 제공하며 sensor configuration, environmental condition, sensing system 간 transfer에 여전히 민감하다.
- Linguistic and acoustic signals: Text 와 audio 는 cross-modal understanding, controllable generation, instruction-conditioned modeling을 위한 의미적·의사소통적 인터페이스를 제공한다.이들은 직접 관측 가능한 신체 상태를 넘어선 행동을 기술하지만 geometry와 motion은 간접적으로 포착하며, 정보성은 semantic specificity, recording quality, transcription accuracy, temporal alignment에 따라 달라진다.
3.2 계산 아키텍처
Section 3.2는 인코딩된 입력이 특정 역량에 맞는 출력으로 변환되는 방식에 따라 human-centric foundation-model architecture를 네 가지 paradigm으로 분류한다. 이러한 paradigm은 단일 단계 매핑부터 순차적·반복적·hybrid computation까지 포괄한다.
- Single-step mapping architectures: Single-step mapping architectures는 하나의 feedforward 계산을 통해 입력을 target representation 또는 structured output으로 변환하며, encoder-only, encoder-decoder, JEPA-style 설계를 포함한다.Encoder-only 모델은 human representation을 생성하고, encoder-decoder 모델은 structured prediction을 생성하며, JEPA-style 모델은 target embedding을 추정한다.
- Sequential factorization architectures: Sequential factorization architectures는 conditional next-token prediction을 사용해 출력을 순서가 있는 시퀀스로 모델링하며, 장거리 의존성(long-range dependencies)을 포착하면서 text, motion, actions [297]을 지원한다.Autoregressive LLM 및 MLLM 아키텍처는 이 정식화를 linguistic symbol 또는 이산화된 human representation에 적용한다.
- Iterative generation architectures: Iterative generation architectures는 초기 상태를 target distribution을 향해 점진적으로 변환하며, diffusion은 reverse denoising을 사용하고 flow-based models [346]은 continuous transport field를 학습한다.Token-by-token generation과 달리 두 방식 모두 출력을 전역적으로 변화하는 상태로 다룬다.
- Hybrid architectures: Hybrid architectures는 sequential, shared-representation, parallel 또는 iterative-feedback composition을 통해 두 가지 이상의 계산 메커니즘을 결합하여 최종 출력을 생성한다.이 메커니즘은 intermediate representation을 결합하는 composition function을 통해 통합될 수 있다.
3.3 최적화 전략
최적화 전략은 학습 중 human-centric capability를 획득하고 추론 중 이를 이끌어내는 방식을 결정한다. 학습에서는 모든 파라미터 또는 일부 파라미터를 변경하는 반면, 추론에서는 학습된 파라미터를 고정한 채 테스트 시점의 계산을 개선하며, 두 단계 모두 결합 가능한 메커니즘을 지원한다.
- 3.3.2 추론: 추론 최적화는 모델 파라미터를 고정하고 semantic augmentation, guided sampling, iterative refinement, preference-based selection, retrieval augmentation을 통해 capability elicitation을 개선한다.이러한 메커니즘은 하나의 추론 과정 안에서 결합될 수 있다.
- 3.3.1 학습: 학습에는 scratch-based training, full-parameter fine-tuning, parameter-efficient tuning, instruction tuning, reward-based tuning, knowledge distillation, task-head tuning의 일곱 가지 메커니즘이 포함되며, 이들은 multi-stage pipeline 전반에서 결합될 수 있다.이러한 메커니즘은 초기화 방식, 업데이트되는 파라미터, 최적화 신호에서 차이를 보인다.
- 3.3.1 학습: 학습 전략은 서로 다른 데이터 및 자원 제약하에서 capability acquisition, adaptation, general knowledge 보존, computational cost, transferability 사이의 균형을 조정한다.Scratch training은 대규모 데이터와 계산을 요구한다. Full fine-tuning은 상당한 domain adaptation을 지원하지만 general capabilities를 약화시킬 수 있으며, parameter-efficient tuning과 task-head tuning은 adaptation cost를 줄이고 feature transfer를 검증할 수 있다.
- 3.3.1 학습: Instruction, reward-based, distillation training은 직접적인 supervised target이 불충분하거나 capability가 아키텍처, modality, pipeline stage에 걸쳐 있을 때 controllability, behavior alignment, capability transfer를 향상시킨다.Instruction tuning은 공유된 language-facing interface를 통해 capability를 드러낸다. Reward-based tuning은 평가된 출력 또는 reinforcement-learning-style objective를 사용한다. Distillation은 inference cost를 줄이면서 출력, 상태 또는 behavior를 전달한다.
- 3.3.2 추론: 추론 메커니즘은 입력을 보강하고, 생성을 유도하며, 출력을 교정하고, 후보 간 선택을 수행하거나, 외부 지식을 검색하여 ambiguity, controllability, consistency, physical validity, quality, instance-level human-centric reasoning을 개선한다.Iterative refinement는 geometry, temporal coherence, interaction feasibility, physical validity를 다룬다. Retrieval augmentation은 파라미터를 변경하지 않고 관련 예시, context 또는 prior를 추가한다.
4 인간 시각적 외양과 공간 기하
이 절에서는 시각적 외양과 공간 기하를 human-centric intelligence의 관찰 가능한 주체 관점으로 규정한다. 외양은 image-space 단서를 포착하고, 기하는 인체를 명시적 구조 또는 렌더링 가능한 구조로 표현한다. foundation-model 시대에 두 영역은 재사용 가능한 human prior와 유연한 인터페이스를 지향하며 발전하고 있다. 외양은 generalist perception, identity understanding, controllable generation을 중심으로, 기하는 structured modeling과 renderable avatar를 중심으로 구성된다.
- Generalist human perception: Human-centric appearance 방법은 task-specific model에서 벗어나, 지각 과제 전반에 재사용 가능한 지식을 전이하는 specialized pretraining과 unified interface를 지향하고 있다.HumanBench, HAP, Sapiens, DAViD, THFM [324], Hulk 및 facial model은 이러한 흐름을 보여준다. 성능 저하 없이 과제 범위를 확장하는 일은 여전히 핵심 과제다.
- Discriminative identity understanding: Identity understanding은 dataset-specific matching에서 벗어나 이질적인 관찰과 query 전반의 recognition으로 이동하고 있지만, 안정적인 identity를 위해서는 지속되는 특성과 변화하는 appearance cue를 분리해야 한다.최근 model은 face-and-body variation, instruction-guided retrieval, multimodal query, structured identity supervision, language-based grounding [108] [109] [64] [140]을 다룬다. Privacy, demographic variation, uncertainty, adaptability는 여전히 중요한 제약이다.
- Controllable human generation: Controllable human generation은 generic synthesis에서 벗어나 identity, anatomy, garment, pose, editing, animation에 대한 compositional control로 발전하고 있다 [44] [52] [12].핵심 난점은 관련 없는 특성을 보존하면서 목표로 한 human property를 편집하는 것이다. 이를 위해서는 여러 의도를 interference 없이 결합하는 재사용 가능한 foundation architecture가 필요하다.
- Spatial geometry: Spatial geometry는 scalable human prior와 transferable semantic interface를 통해 명시적 인체 구조 또는 렌더링 가능한 avatar asset을 모델링함으로써 human-centric intelligence를 image-space appearance 너머로 확장한다.이 subsection은 structured geometry modeling과 renderable avatar modeling을 구분한다. 후자는 geometry와 appearance를 렌더링 가능한 asset으로 통합한다.
5 인간 운동학적 동역학과 상호작용 모델링
이 절에서는 인간 중심 지능을 동적 행위자의 관점에서 바라보며, 시간에 따른 신체 변화와 객체·환경·사람 간 관계를 다루는 운동학적 동역학과 상호작용 모델링을 제시한다. 운동학적 동역학은 확장 가능한 동작 모델링과 human video animation으로 구성하고, 상호작용 모델링은 human-object, human-scene, social interaction으로 구분한다.
- 절 개요: 운동학적 동역학과 상호작용 모델링은 인간을 객체·환경·다른 사람과의 관계에 의해 시간적 신체 변화가 형성되는 동적 행위자로 규정한다.이 분야는 고립된 동작 생성기와 상호작용별 파이프라인을 넘어 재사용 가능한 temporal prior, multimodal interface, 더욱 일반적인 모델로 나아가고 있다.
- Human video animation: Human video animation은 audio-, pose-, motion-, appearance-, multimodal conditioning을 통해 사실적인 video에서 신체 동역학을 구현하며, video generation과 구조화된 동작을 점차 통합한다.대규모 video backbone은 사실성과 유연성을 높이지만, driving과 generated body의 정밀한 대응, 장기 일관성, identity consistency, 해부학적 타당성, 미세한 타이밍은 여전히 해결되지 않은 문제다.
- 확장 가능한 동작 모델링: 확장 가능한 동작 모델링은 움직임을 temporal signal로 다루며, 재사용 가능한 temporal 및 linguistic interface, motion-language unification, multimodal observation, tool- 또는 retrieval-assisted control을 발전시킨다.대표적인 시스템으로는 transferable masked-sequence feature를 위한 MotionBERT, discrete-token captioning과 generation을 위한 MotionGPT [133], planning과 task decomposition을 위한 AvatarGPT, shared motion understanding과 generation을 위한 MG-MotionLLM 이 있다.
- 확장 가능한 동작 모델링: Motion-native foundation model은 data, architecture, task 전반의 scaling과 transfer를 점차 연구하고 있지만, 현재의 scaling은 physical validity, temporal causality, 미세한 인간 변이를 부분적으로만 포착한다.예로는 Being-M0 [332], ScaMo, Go-to-Zero [70], Kimodo [283], HY-Motion [346]이 있으며, 핵심 과제는 서로 다른 objective와 observation domain에서 continuous kinematics를 보존하는 것이다.
- 상호작용 모델링: 상호작용 모델링은 open-ended semantic knowledge, multimodal reasoning, 재사용 가능한 generative prior를 통해 국소적인 human-object 관계에서 human-scene compatibility와 social coordination으로 확장된다.이 체계는 인간을 외부 개체에 의해 행동이 공동으로 형성되는 관계적 행위자로 다룬다.
6 인간 월드 시뮬레이션과 Embodied Agency
이 절은 월드 시뮬레이션과 Embodied Agency를 human-centric intelligence의 상황적 에이전트 수준으로 규정하며, 지각적 월드 생성과 실행 가능한 계획 및 물리적 행동을 구분한다. Foundation model 접근법을 검토하는 한편, 유용한 시스템에는 지속적이고 인과적으로 반응하며 제어 가능한 월드와 물리적으로 근거된 agency가 필요함을 강조한다.
- 월드 시뮬레이션: 월드 시뮬레이션은 인간 활동을 중심으로 지각적 미래를 모델링하는 human-centered world generation과, 의사결정 및 embodied action을 위해 행동에 의존하는 전이를 예측하는 actionable world planning을 구분한다.이 구분은 Fig. 10에 제시되어 있으며, 시각적으로 그럴듯한 시뮬레이션과 출력 또는 내부 변수가 행동을 지원하는 예측 모델을 나눈다.
- 월드 시뮬레이션: Foundation video model은 인간의 motion, viewpoint, goals, scene structure, full-body trajectories를 점차 더 지속적인 first-person simulations로 확장하지만, 그 출력은 여전히 실행 가능하기보다는 지각적이다.예로는 Generated Reality [361], Hand2World [339], AnchorWorld [199], EgoForge [293], EgoControl, EgoX, EgoWorld [266]가 있다.
- 월드 시뮬레이션: Actionable world model은 pose, manipulation 또는 dexterous actions를 조건으로 미래 예측을 수행하며, visual forecasting을 action generation, interaction learning, closed-loop planning과 점차 결합한다.대표적인 시스템으로 PEVA [8], LOME [87], EgoHOI [175], HandWorld, DexWM [91], EgoAgent [43], EgoSim [106], EgoExo-WM [315]가 있다.
- 월드 시뮬레이션: 지각 정확도만으로는 decision utility를 보장하지 않는다: actionable world model에는 신뢰할 수 있는 intervention response, 결과를 초래하는 변화와 우연한 변동의 분리, calibrated uncertainty가 필요하다.Generative fidelity와 planning effectiveness의 관계는 human-centered world modeling에서 여전히 핵심 과제로 남아 있다.
- Embodied Agency: Embodied Agency는 generalist humanoid control과 human-to-agent skill transfer를 구분하며, foundation model을 semantic engine으로 활용하는 동시에 재사용 가능한 motor prior, specialized controller, scalable behavioral pretraining을 함께 사용한다.대표적인 방향으로 VGHuman [392], HOI-HLI, BiBo [132], VLM-RMD [60], MaskedMimic, InterMimic, TokenHSI, SCRIPT [414], SONIC [240], Humanoid-GPT [275]가 있다.
7 데이터셋, 벤치마크, 메트릭
이 절에서는 human-centric model의 개발과 평가를 뒷받침하는 대표 데이터셋, 벤치마크, 널리 사용되는 메트릭을 검토하고, 현재 평가 관행의 한계를 짚는다.
- 7 데이터셋, 벤치마크, 메트릭: 데이터셋과 벤치마크는 human-centric model이 무엇을 학습하고 그 역량을 어떻게 평가받는지를 규정한다.이 절에서는 대표 리소스를 검토한 뒤 널리 사용되는 메트릭을 체계적으로 정리한다.
- 7 데이터셋, 벤치마크, 메트릭: 체계적으로 정리된 메트릭은 현재 평가 관행을 명확히 보여주고, 그러한 관행이 여전히 불완전한 지점을 드러낸다.
7.1 데이터셋과 벤치마크
데이터셋은 관측과 주석을 제공하고, 벤치마크는 인간 중심 모델의 역량을 비교할 수 있도록 과제와 프로토콜을 표준화한다. 리소스는 인간 주체, 인간 동역학, 인간 상호작용, 인간 체화라는 네 가지 taxonomy-aligned 그룹으로 구성되며, 점차 풍부해지는 지각적·시간적·관계적·체화 맥락을 포괄한다.
- 구성: 리소스 taxonomy는 데이터셋과 벤치마크를 인간 주체, 인간 동역학, 인간 상호작용, 인간 체화로 구분해, 이질적인 리소스를 인간 맥락의 발전 과정과 연결한다.데이터셋은 관측과 주석을 통해 모델 개발을 지원하는 반면, 벤치마크는 정의된 과제와 프로토콜을 통해 표준화된 비교를 가능하게 한다.
- 인간 주체: 인간 주체 리소스는 시각 및 다중감각 관측에서 identity correspondence와 renderable 3D or 4D human geometry로 발전한다.이들은 외형 모델링, 가림 또는 프라이버시 제약하의 multimodal sensing, 개인 수준 매칭, 재구성, animation, relighting, controllable synthesis를 지원한다.
- 인간 동역학: 인간 동역학 리소스는 시간적으로 구조화되거나 조건화된 데이터를 통해 video generation과 animation, 행동 이해, 운동학적 동작, gait, 스포츠 분석, virtual try-on을 다룬다.이들은 세밀한 motion reasoning, motion-language supervision, 변화하는 관측 조건에서의 인식, 운동선수 피드백, 정체성을 보존하는 의류 또는 신발 전이를 점차 지원한다.
- 인간 상호작용: 인간 상호작용 리소스는 절차적, 물리적, 환경적, 사회적 맥락에서의 행동을 표현하며, egocentric activity, human-object interaction, human-scene interaction, social interaction으로 구성된다.Egocentric procedural 데이터셋은 행위자의 관점에서 단계 인식, 의도 예측, 오류 탐지, 지원을 가능하게 하며, 점점 풍부해지는 손, 지원, 조밀한 3D 주석을 포함한다.
7.2 Metrics
이 절은 인간 중심 평가를 재구성 중심, 의미 중심, 생성 중심, 상호작용 중심, 효율성 중심의 다섯 metric family로 구성한다. 이는 이질적인 출력이 geometry, semantics, generated content, physical interactions, executable behavior를 포괄하기 때문이다. 또한 출력 일치도와 의미 보존에서 generative quality와 contextual validity에 이르기까지 점점 확장되는 평가 대상을 기준으로 이러한 metrics를 체계화한다.
- Metric Taxonomy: 인간 중심 평가는 이질적인 출력과 평가 대상을 포괄하기 위해 재구성 중심, 의미 중심, 생성 중심, 상호작용 중심, 효율성 중심 metrics로 분류된다.이 taxonomy는 geometric structures, semantic predictions, generated content, physical interactions, executable behavior를 아우른다.
- Reconstruction-Centric Metrics: 재구성 중심 metrics는 articulated human state, surface geometry, image 또는 rendering fidelity 전반에서 reference observations와의 일치도를 측정한다.관절 정확도에는 MPJPE와 PA-MPJPE, dense meshes에는 PVE/MPVPE, surfaces에는 Chamfer Distance와 Surface F-score, visual fidelity에는 PSNR, SSIM, LPIPS 등이 사용된다.
- Semantic-Centric Metrics: 의미 중심 metrics는 recognition, retrieval 및 verification, semantic alignment를 통해 task-relevant, identity-related, cross-modal meaning을 평가한다.여기에는 accuracy, precision, recall, F1, AP/mAP, Rank-k, CMC, mINP, R-Precision, Recall@K, CLIP- 및 DINO 기반 similarities, language-generation metrics가 포함된다.
- Generative-Centric Metrics: 생성 출력에는 fidelity, diversity 및 coverage, temporal dynamics, audio-visual synchronization, identity 또는 appearance consistency를 평가하는 상호보완적 metrics가 필요하다.이 차원들은 FID와 FVD 같은 distributional measures, diversity 및 coverage measures, trajectory 및 smoothness diagnostics, synchronization scores, identity 또는 garment consistency measures를 사용한다.
- Interaction-Centric Metrics: 상호작용 중심 metrics는 contact와 collision, physical plausibility, embodied task performance, human preference 또는 subjective quality 전반에서 예측되거나 생성된 behavior가 유효하게 유지되는지를 평가한다.이 구성은 local geometric validity에서 physical execution, task completion, perceived naturalness로 단계적으로 확장된다.
8 개방형 과제와 향후 방향
이 분야는 범용 human-centric 시스템을 향해 나아가고 있지만, 개별 과제만 확장하는 것으로는 데이터, 통합 표현, 물리적 grounding, world modeling, 평가, 배포의 과제를 해결할 수 없다. 여섯 가지 방향은 확장 가능하고 전이 가능하며 신뢰할 수 있고 물리적으로 grounded된 효율적 human-centric intelligence를 지향한다.
- 데이터 확장: 실제 인간 데이터의 제한성은 automatic annotation, 낮은 marginal cost, 어려운 real-world 사례에 대한 targeted coverage를 통해 가능해지는 합성 데이터 확장을 촉진한다.데이터 수집에는 specialized capture systems, 광범위한 annotation, personal information에 대한 신중한 처리가 필요하다.
- 데이터 확장: 향후 연구는 데이터 quantity와 quality를 분리하고, 실제 및 underrepresented population을 대상으로 real-synthetic mixture를 검증하며, adaptation 과정에서 provenance와 consent를 보존해야 한다 [9].이러한 protocol은 synthetic data가 언제 실질적인 transfer를 제공하는지 명확히 할 수 있다.
- 통합 Foundation Model: 일반적인 human-centric foundation model은 full human context spectrum 전반에 걸친 unified representation learning을 통해 modality- 또는 task-restricted system을 대체해야 한다 [12] [324].현재의 unified model은 대체로 제한된 task group이나 특정 modality만 다룬다.
- Physical Grounding: 차세대 representation은 physical information을 중심적으로 모델링해야 한다. visual 및 linguistic regularity만으로도 이를 가능하게 하는 physical law를 포착하지 않은 채 그럴듯한 behavior를 지원할 수 있기 때문이다 [35] [92] [122].Physical grounding은 post-prediction 또는 post-generation constraint로만 적용되어서는 안 된다.
- World Model: Human-centric world model은 agent action과 environmental response를 공동으로 예측해야 한다. 시각적으로 그럴듯한 future만으로는 causal action-consequence modeling이 성립하지 않기 때문이다 [293] [339] [361].Human video는 확장 가능한 experience를 제공하지만, embodiment gap은 관찰된 behavior에서 직접 transfer하는 데 한계를 둔다.
- 평가: 평가는 perception, generation, interaction 전반의 benchmark를 연결해야 한다. 분리된 task performance만으로는 cross-context transfer나 capability composition을 입증할 수 없기 때문이다 [28] [323].핵심 한계는 metric 자체의 부재가 아니라 기존 metric을 연결하는 protocol의 부재다.
- 배포: 효율적인 deployment에는 monolithic model의 대안이 필요하며, agentic, tool-augmented coordination을 통해 training cost, inference latency, privacy exposure, maintenance complexity를 해결해야 한다.이러한 system은 foundation model, specialized human model, external perception tool을 조정할 수 있다.
9 결론
본 survey는 six-level human context taxonomy와 methodological foundations를 바탕으로 foundation-model era의 human-centric intelligence를 전 범위에 걸쳐 체계적으로 제시한다.
- 결론: 본 survey는 foundation-model era의 human-centric intelligence를 전 범위에 걸쳐 체계적으로 제시한다.
- 결론: human context taxonomy는 visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, embodied agency를 연결한다.
- 결론: 이 taxonomy는 인간을 observable subjects, dynamic actors, situated agents로 규정하는 동시에 field의 methodological foundations와 발전 성과에 대한 review를 이끈다.