Source-linked AI summary

A Comprehensive Review of Generative Physical Artificial Intelligence

Satyam Gaba, Krutiksinh Rana, Siva Sai, Vinay Chamola, Dusit Niyato

arXiv:2609.18111v1cs.ROcs.AIcs.CLcs.CVcs.LG

TL;DR

GPAI research lacks an integrated framework connecting model types across the perception–cognition–actuation pipeline while addressing safety, data, and deployment challenges. This survey organizes five model classes by control-pipeline function, finds complementary specialization and integration, and reports improved generalization in representative robotic applications.

  • Problem

    Existing surveys overlook the full range of GPAI model types and do not systematically connect generative techniques across the perception–cognition–actuation pipeline or cross-cutting deployment challenges.

  • Method

    The survey develops a taxonomy mapping RFMs, VLAs, WFMs, LBMs, and Diffusion Policies to robotic control-pipeline functions and synthesizes their architectures, applications, and limitations.

  • Results

    RT-2 achieves a 62% success rate on unseen tasks versus 32% for RT-1, demonstrating improved generalization and semantic reasoning in robotic control.

  • Takeaways & Limitations

    The taxonomy shows that integrated and specialized models remain complementary, with capabilities increasingly shaped by data-centric learning and convergence across model classes.

  • Takeaways & Limitations

    Earlier robotic systems were constrained to narrowly defined, pre-programmed tasks and could not adapt to unexpected scenarios.

Abstract

from arXiv · show

The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI). These agentic AI systems autonomously perceive, reason, and act in complex real-world situations. This survey comprehensively analyzes GPAI systems, focusing on their architectural foundations, current applications, and key limitations. We introduce a taxonomy of five distinct approaches: Robot Foundation Models (RFMs) for cross-platform skill transfer; Vision-Language Action (VLA) models for end-to-end multi-modal perception and control; Large Behavior Models (LBMs) for human-like movement generation; Diffusion Policy Models (DPMs) for diffusion model-based temporally coherent action generation; and World Foundation Models (WFMs) for physics-compliant simulation and data generation. We examine how these approaches complement each other: WFMs generate training data for VLAs and DPMs, RFMs enable cross-platform deployment of learned policies, while LBMs provide motion priors for natural behavior. Through examples across autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems, we identify significant performance improvements and summarize promising research directions in data-efficient learning, sim-to-real transfer, edge-compatible architectures, and safety frameworks. These insights advance embodied AI for IoT-connected environments where intelligent agents interact with networked sensors, actuators, and edge devices.

I. INTRODUCTION

GPAI extends foundation-model capabilities into closed-loop physical control, addressing fragmented survey coverage while introducing an integrated view of architectures, applications, and deployment challenges.

  • I. INTRODUCTION: GPAI applies large-scale generative models to synthesize actions, trajectories, and environment predictions for autonomous physical systems.It emphasizes zero-shot generalization and iterative policy generation rather than task-specific direct observation-to-action mappings.
  • I. INTRODUCTION: The survey frames GPAI as an evolution from narrow, pre-programmed systems toward foundation-model-based autonomy across tasks, embodiments, and environments.Early systems used deterministic rules and scripted motions that failed under changed part positions, tool wear, or unexpected obstacles.
  • I. INTRODUCTION: Unlike rigid hierarchical or reactive robotics, modern GPAI combines generative reasoning with perception and control to improve contextual adaptation in unstructured environments.Hybrid architectures historically combined planning and reactive execution; GPAI enhances these components with LLMs and VLMs.
  • I. INTRODUCTION: Industrial adoption is growing, but deployment depends on safety certification, regulation, integration costs, and workforce considerations.Reported examples include more than 750,000 AI-driven robots at Amazon and efficiency gains of up to 25% in some fulfillment centers.
  • I. INTRODUCTION: The survey addresses fragmented prior coverage by connecting GPAI models across the perception–cognition–actuation pipeline and incorporating applications, benchmarks, safety, sim-to-real transfer, data scarcity, and ethics.Prior surveys were dispersed across manipulation, physics-aware perception, foundation models, and embodied AI.

D. Research Contribution and Scope

The survey establishes a comprehensive scope spanning GPAI architectures, components, applications, benchmarks, and deployment risks, organized around an end-to-end embodied system.

  • D. Research Contribution and Scope: The survey contributes a five-paradigm taxonomy—RFMs, VLAs, LBMs, DPMs, and WFMs—covering perception–action integration, policy generation, and world simulation.It also presents component views, evaluation benchmarks, applications, and limitations across robotics and autonomous domains.
  • D. Research Contribution and Scope: The survey covers applications in autonomous vehicles, manufacturing, healthcare, humanoids and HRI, logistics, digital infrastructure, and consumer or research settings.Its scope includes both capabilities and persistent gaps across these domains.
  • D. Research Contribution and Scope: The paper is organized around GPAI components, taxonomy, applications, challenges and risks, and future directions.Section II covers the closed-loop system; Section III develops the five architectures; Section IV surveys applications; Section V discusses challenges and remedies.
  • D. Research Contribution and Scope: GPAI embodiments integrate cognition, perception, actuation, and data into an autonomous architecture that continuously interacts with real-world environments.Feedback from actions returns to perception and data components for evaluation and refinement, while simulation supports training and edge-case coverage.
  • D. Research Contribution and Scope: The end-to-end architecture transforms multimodal inputs and high-level intentions into trajectories, joint commands, and actuator signals under mechanical and workspace constraints.Perception uses sensors such as cameras, IMUs, LiDAR, and RADAR, while controllers provide position, velocity, and torque commands.

D. Perception

GPAI perception combines multimodal sensing, demonstrations, simulation, and learning mechanisms to support robust understanding and action in changing physical environments.

  • D. Perception: Perception supports closed-loop adaptation by integrating camera, IMU, LiDAR, and RADAR streams with cognition and actuation.This integration is presented as a basis for handling uncertainty across diverse tasks and conditions.
  • D. Perception: GPAI perception draws on RGB-D, audio, force/torque, tactile, proprioceptive, environmental, demonstration, synthetic, and multimodal data.These sources connect semantic information with physical interaction and encode expert state–action knowledge.
  • D. Perception: Simulation and digital twins provide scalable, risk-free scenarios covering rare events, extreme conditions, manipulation, navigation, and human–robot interaction.Physics simulators generate photorealistic environments with dynamics that support extensive training before physical deployment.
  • D. Perception: Large-scale pretraining across tasks, embodiments, environments, lighting, and weather improves generalization, while data fidelity remains critical for unpredictable settings.Structured, unstructured, and real-time-stream data serve distinct purposes in skill learning and environmental understanding.
  • D. Perception: GPAI uses reinforcement, hierarchical, supervised, unsupervised, self-supervised, and human-feedback learning to acquire motor skills and representations.Hierarchical reinforcement learning separates strategic meta-control from fine-grained execution, while RLAIF scales preference optimization using auxiliary-model comparisons.

G. Infrastructure and Evaluation Frameworks

GPAI infrastructure organizes model families by their roles in control, policy generation, and data synthesis, while RFMs use multimodal architectures to generalize across embodiments.

  • G. Infrastructure and Evaluation Frameworks: RFMs and VLAs span perception, planning, and action, whereas WFMs synthesize predictive simulations, LBMs generate human-like interaction, and Diffusion Policies provide low-level manipulation control.The taxonomy therefore links model design to generality, specialization, data demands, adaptability, and applications.
  • G. Infrastructure and Evaluation Frameworks: RFMs process action, language, sensor, and visual inputs through specialized encoders, cross-modal fusion, and an action decoder that outputs commands across robotic embodiments.This architecture supports unified representations for transferring capabilities across tasks and platforms.
  • G. Infrastructure and Evaluation Frameworks: General-purpose models integrate perception, planning, and control across heterogeneous platforms, but require substantially greater computational resources and large-scale datasets than task-specific models.Task-specific systems can be accurate in constrained repetitive settings but often require reconfiguration or retraining for new tasks.
  • G. Infrastructure and Evaluation Frameworks: Representative RFM designs include language-centric embodied transformers such as PaLM-E and unified token-sequence modeling such as Gato.Gato tokenizes heterogeneous inputs and outputs into a shared vocabulary processed by one transformer.

B. Vision Language Action (VLA) Models

VLA models unify visual perception, language understanding, and robotic action generation through multimodal architectures that directly map inputs to motor commands. Implementations such as RT-1, RT-2, OpenVLA, and Gemini Robotics VLA demonstrate scalable generalization, efficient control, and increasingly dexterous behavior.

  • VLA Models: VLAs combine multimodal reasoning components with transformer or diffusion-based action modules for end-to-end sensorimotor control.Their specialized visual encoders, semantic projectors, and language models translate multimodal inputs directly into executable motor commands.
  • Multimodal Processing Pipeline: The VLA pipeline separately encodes images, language, and proprioceptive states, fuses them into contextualized embeddings, and decodes action tokens for robot control.TokenLearner reduces visual representations from 81 tokens to 8 salient tokens, lowering downstream computational cost while maintaining performance.
  • RT-2: Actions as Language: RT-2 represents robotic actions as language, enabling semantic knowledge from pretrained vision-language models to transfer into executable behaviors.The approach is reported to improve generalization to novel tasks and environments, object-action reasoning, and multimodal knowledge transfer.
  • Other Contemporary VLA Implementations: OpenVLA surpasses larger closed-source RT-2-X models across 29 benchmarks while using seven times fewer parameters.It combines DINOv2 and SigLIP visual features with Llama 2 and discretizes seven-degree-of-freedom actions into 256 bins.
  • Other Contemporary VLA Implementations: Gemini Robotics VLA extends VLA control to coordinated dual-arm manipulation through integration with the ALOHA teleoperation platform.This highlights embodied systems capable of more dexterous, human-like physical interaction than traditional single-arm manipulation.

C. Large Behavior Models (LBMs)

Large Behavior Models learn and generate adaptive, human-like physical action sequences from multimodal data, using reinforcement learning and feedback on behavioral realism and physical coherence. Diffusion Policy Models complement this approach by iteratively denoising noisy action trajectories, yielding temporally coherent and multimodal control.

  • Large Behavior Models: LBMs learn complex physical behaviors from multimodal data and use reinforcement learning to adapt actions to feedback, uncertainty, and changing environments.This adaptability supports robotics, embodied assistance, and interactive coaching applications requiring physical engagement and situational awareness.
  • Iterative Action Refinement Framework: LBM pretraining combines expert demonstrations with authenticity feedback from discriminators and plausibility feedback from physics-coherence evaluations.These rewards refine actor policies, while inference networks generate real-time actions from state inputs and latent variables.
  • Large Behavior Models: Meta Motivo demonstrates zero-shot whole-body humanoid control without task-specific training through coordinated networks based on FBCPR.Its architecture converts state observations into humanoid actions while targeting human-like motion generation.
  • Model Architecture and Advantages: Diffusion Policy architectures encode multimodal observations and iteratively predict noise to transform Gaussian noise into executable action sequences.The process uses stacked observations, FiLM conditioning, and cross-attention to preserve temporal context and adapt to visual and proprioceptive conditions.
  • Model Architecture and Advantages: Diffusion Policy models generate temporally coherent action sequences extending over 4–8 future steps, improving stability for continuous control.Receding-horizon execution performs initial predicted actions and frequently recalculates plans to limit error accumulation in dynamic environments.
  • Model Architecture and Advantages: Diffusion Policy models report a 47% average success-rate improvement across 12 robotic tasks in 4 benchmark suites.Their stochastic denoising process represents multiple valid solutions and avoids mode collapse associated with deterministic behavior cloning.

3) Training and Evaluation:

Training and evaluation for physical AI combine grounded robotic data, predictive world models, and closed-loop planning with task- and sequence-level metrics. Diffusion policies increasingly support spatial generalization, stable fine-tuning, and temporally coherent control under data and reward constraints.

  • Diffusion Policy Models: Diffusion-policy evaluation emphasizes task success, multimodal action coverage, temporal consistency, and systematic testing across precision tasks, environments, and robot configurations.RoboMimic supplies diverse teleoperated manipulation trajectories, while Push-T tests exact geometric alignment.
  • Diffusion Policy Models: Diffusion policies increasingly combine 3D spatial generalization, reinforcement-learning fine-tuning, and structured denoising to improve robustness under sparse rewards and limited demonstrations.DP3 uses point clouds for spatial generalization, while DPPO supports stable fine-tuning; task-horizon noise schedules and structured latent spaces address long-horizon consistency and physical constraints.
  • World Foundation Models: WFMs predict future environment representations from tokenized observations and actions, enabling simulated rollouts for counterfactual planning, validation, and synthetic data generation.World predictions may be images, videos, latent trajectories, or sensor outputs, but longer horizons compound prediction errors.
  • World Foundation Models: WFM training uses curated external logs and robot interactions, synchronizing multimodal streams and tokenizing video, proprioception, LiDAR, and inertial data.The data pipeline begins with episode segmentation and quality filtering, with optional task-relevant annotation.
  • World Foundation Models: Closed-loop WFM control samples candidate action sequences, scores their simulated outcomes with costs, rewards, and safety penalties, then executes the selected sequence through receding-horizon replanning.A policy can be distilled from planner-selected actions or improved through model-based reinforcement learning, while new execution data feeds continued refinement.

1) Environmental Simulation and Data Generation:

Environmental simulation combines learned world models with digital twins and physics-informed constraints to generate controllable, realistic rollouts and downstream training data. Evaluation spans perceptual quality, controllability, dynamics, physical plausibility, and embodied utility, while continual-learning updates require safety validation.

  • Environmental Simulation: Hybrid pipelines combine digital twins’ interpretable, constraint-driven simulation with WFMs’ learned rollouts for complex dynamics and observation statistics that are difficult to model analytically.Digital twins remain synchronized virtual replicas, while WFMs provide data-driven predictive capabilities.
  • Physics-Informed Simulation: Physics-informed learning embeds governing equations and conservation laws into WFM losses or latent spaces to improve physical realism and reduce sim-to-real discrepancies.The paper highlights thermodynamics and material properties as especially important in intelligent manufacturing.
  • Synthetic Data Generation: WFMs generate controllable synthetic scenario families from varied initial states and actions, amplifying rare events across lighting, weather, and scene conditions for perception and control.Examples include near-collision driving scenarios and manipulation failures that are difficult to collect extensively in the physical world.
  • Continual Learning: Continual-learning pipelines use deployment data for policy distillation, model-based reinforcement learning, and calibration, but require validation to prevent safety-critical regressions.Periodic grounding on real interactions is used to mitigate model bias and drift.
  • Evaluation: WFM evaluation measures generation quality, temporal stability, controllability, physical plausibility, and downstream decision-making utility through perceptual, physics-oriented, task-oriented, and human assessments.WorldScore, WorldModelBench, and WorldSimBench evaluate these dimensions using structured scenarios, human labels, physics adherence, and video-to-action utility.

F. Taxonomy Synthesis and Discussion

The taxonomy maps five GPAI model classes to complementary stages of embodied control, balancing general-purpose integration with specialized capabilities. Applications span autonomous driving, industrial robotics, healthcare, humanoids, assistive technologies, manufacturing, and digital twins.

  • Taxonomy Synthesis: The taxonomy identifies a generality–specialization tradeoff: RFMs and VLAs span perception, planning, and action, while WFMs, LBMs, and DPMs provide specialized capabilities.WFMs emphasize simulation and data generation, LBMs human-like interaction, and diffusion policies robust low-level manipulation control.
  • Taxonomy Synthesis: The field is shifting from modular task-specific systems toward integrated end-to-end architectures trained on large, diverse, and increasingly data-centric inputs.The survey describes emerging convergence in which diffusion and behavior-model strengths integrate with generalist RFMs and VLAs.
  • Applications: GPAI applications connect perception, reasoning, and physical action across autonomous driving, robotics, healthcare, humanoids, assistive technologies, manufacturing, and digital twins.Figure 8 summarizes these application domains.
  • Autonomous Vehicles: In autonomous driving, WFMs generate realistic and diverse scenarios for training and validation, while GAIA-1 produces controllable rollouts from tokenized video, text, and action inputs.Tesla’s fleet-learning approach instead continuously collects driving situations, retrains models, and distributes updates over the air.
  • Industrial Robotics and Manufacturing: Industrial GPAI enables zero-shot handling of novel objects and tasks, while diffusion policies generate temporally coherent sequences for long-horizon manipulation.A UR10 deployment palletized 7 kg packages at six per minute across more than 350 product variants without safety barriers.

C. Healthcare and Medical Robotics

GPAI is being applied across healthcare, human-robot interaction, logistics, digital twins, and service robotics to support multimodal understanding, adaptive control, and operational optimization. Reported examples include clinical systems, faster warehouse fulfillment, and virtual replicas for planning and synthetic data generation.

  • Healthcare and Medical Robotics: 98.5% surgical success in FDA trials for Medtronic’s Hugo platform exceeded the 85% benchmark, while da Vinci surpassed 14 million minimally invasive procedures.Hugo supports procedure planning, real-time 3D imaging, and instrument guidance; da Vinci provides precision, dexterity, and control across specialties.
  • Humanoid Robotics and Human-Robot Interaction: NVIDIA’s 2.2B-parameter Groot N1 combines vision, language, proprioception, and diffusion-based action refinement with 64ms inference.Its VLM operates at 10Hz while the Diffusion Transformer generates actions at 120Hz, supporting smooth and reliable movement.
  • Humanoid Robotics and Human-Robot Interaction: 62% success on unseen tasks for RT-2 compared with RT-1’s 32% demonstrates improved generalization and semantic reasoning in robotic control.RT-2 transfers knowledge from large-scale web data to recognize and act in novel situations without explicit task-specific training.
  • Logistics, Warehousing, and Supply Chain Management: Amazon has deployed more than 750,000 robots, while its Sequoia platform improves fulfillment-center productivity by approximately 25%.GPAI supports dynamic task allocation, path planning, scheduling, and manipulation of diverse objects in unstructured warehouse settings.
  • Digital Infrastructure and Smart Systems: BMW’s digital twins span more than 30 production sites and are projected to reduce production-planning costs by up to 30%.These virtual replicas support predictive maintenance, process optimization, scenario planning, and synthetic-data generation; Regensburg monitors about 1,400 vehicles daily.
  • Digital Infrastructure and Smart Systems: Digital twins generate high-fidelity virtual replicas that support predictive maintenance, process optimization, scenario planning, and GPAI training and evaluation.They can reduce financial and operational risks associated with real-world deployment.

B. Bias and Ethical Concerns

GPAI faces ethical and deployment risks from biased data, computational constraints, reasoning failures, integration complexity, and limited transparency. These risks can produce unsafe or discriminatory physical behavior and complicate reliable real-world operation.

  • Bias and Ethical Concerns: Biased training data can produce brittle, prejudiced, unsafe, discriminatory physical actions, and legal or regulatory liabilities.The paper frames bias mitigation as an ethical responsibility, a regulatory requirement, and a precondition for trustworthy GPAI.
  • Generalization and Reasoning: Poor contextual reasoning can impair understanding of physical actions, causing models to misinterpret relations or fail on logical outcomes.The paper identifies advanced reasoning models and multimodal integration as directions for addressing this limitation.
  • Safety and Robustness: Integrating perception, reasoning, and control modules is complex, and increasing complexity can reduce reliability when scaling from laboratories to physical systems.Physical failures can cause harm to humans or property and may involve critical failure modes absent from digital systems.
  • Energy Efficiency: Training demands massive computational power with a significant carbon footprint, motivating efficient algorithms and green computing.This environmental constraint accompanies the broader deployment requirements for GPAI systems.
  • Transparency and Explainability: Black-box decision processes make auditing and error diagnosis difficult, increasing the need for explainable and interpretable models.The paper also identifies rigorous safety checks, monitoring, and fail-safes as emerging safeguards.
  • Hardware and Real-Time Constraints: Large GPAI models can conflict with real-time edge requirements because memory, compute, energy, and iterative denoising constrain feasible deployment.Compression and optimization may reduce accuracy, generalizability, and reliability, while diffusion policies can conflict with sub-50ms manipulation response times.

D. Simulation-to-Real Transfer and Uncertainty

Simulation-to-real transfer remains difficult because simulations simplify physical variation, contact dynamics, actuator imperfections, and sensor noise. Models trained only in clean simulations can fail under real-world uncertainty, although domain randomization, transfer learning, and reality-gap refinement are narrowing the divide.

  • Simulation-to-Real Transfer and Uncertainty: LBM performance often degrades under unmodeled contact scenarios and hardware imperfections, with limited real-world validation beyond basic locomotion.This makes simulation-to-real transfer a foundational reliability and safety requirement rather than only a performance concern.
  • Simulation-to-Real Transfer and Uncertainty: Simulation can omit friction, mass variation, nonlinear contact dynamics, actuator imprecision, and pervasive sensor noise.Models trained on clean environments may consequently overfit simulation and produce behaviors that collapse under real-world uncertainty.
  • Simulation-to-Real Transfer and Uncertainty: Domain randomization, sim-to-real transfer learning, and iterative reality-gap refinement are being used to narrow the simulation-to-real divide.Differentiable simulation and physics-informed training can enhance fidelity, but simulations still simplify physical-world complexity.

E. Generalization, Prompt Sensitivity, and Reasoning Challenges

GPAI generalization remains constrained by embodiment mismatch, reasoning and prompt sensitivity, system integration complexity, and unresolved safety and transparency requirements. These challenges motivate modular architectures, runtime safeguards, and broader adaptation strategies.

  • E. Generalization, Prompt Sensitivity, and Reasoning Challenges: Embodiment mismatch across actuator dynamics, joint limits, and sensing configurations can make Robot Foundation Model policies unpredictable on new robots.Continual fine-tuning may also cause catastrophic forgetting of previously learned skills.
  • E. Generalization, Prompt Sensitivity, and Reasoning Challenges: Reasoning failures can misinterpret object relationships or long-term consequences, cascading into unsafe behavior, while slight prompt variations can produce inconsistent or hazardous outputs.Open-ended generalization additionally requires continual adaptation, novel affordance inference, and skill acquisition without explicit retraining.
  • F. Scalability and System Complexity: Integrating heterogeneous perception, reasoning, planning, and control modules remains a central scaling barrier because failures in one layer can destabilize the entire system.The reliability burden grows exponentially with task complexity and environmental variability.
  • G. Robustness, Safety, and Failure Modes: Physical deployment requires near-perfect reliability because reasoning, integration, or control errors can cause cascading failures that harm people, infrastructure, or the environment.The paper calls for safety evaluation, continuous real-world monitoring, and fail-safe mechanisms across the model lifecycle.
  • G. Robustness, Safety, and Failure Modes: Fixed action-token vocabularies in vision-language-action models improve training manageability but reduce control resolution for finely calibrated manipulation.For example, discretizing actions into 256 bins per dimension can introduce motion discontinuities that hinder smooth execution.
  • G. Robustness, Safety, and Failure Modes: Diffusion-based controllers lack widely adopted formal stability and safety certification methods, leaving their disturbance behavior primarily validated empirically.Runtime Control Barrier Function filters provide partial mitigation by constraining policies to defined safe operating regions, but composition with stochastic foundation-model outputs remains open.
  • H. Transparency and Explainability: Black-box GPAI models complicate error diagnosis, accountability, auditing, and user confidence, especially in safety-critical domains requiring transparent decisions.Interpretability is therefore presented as a regulatory, ethical, and trust requirement rather than merely a technical aspiration.

I. Energy Efficiency and Environmental Impact

The paper identifies computational cost and environmental impact as concerns for scaling GPAI, while outlining efficiency, simulation, safety, and deployment directions. Future progress emphasizes more data-efficient, capable, robust, and adaptable systems.

  • I. Energy Efficiency and Environmental Impact: Large-scale GPAI training consumes substantial computational resources, causing high energy use and notable carbon emissions as model sizes and workloads scale.The paper frames performance gains against ecological responsibility and carbon efficiency.
  • I. Energy Efficiency and Environmental Impact: Data-efficient learning, cross-embodiment foundation models, and computational efficiency for edge devices are identified as critical research directions for practical deployment.The stated goals include reducing annotated-data dependence, enabling transfer across embodiments, and supporting real-time edge operation.
  • I. Energy Efficiency and Environmental Impact: Self-supervised, transfer, and few-shot learning are proposed to support effective learning from smaller, less curated datasets and agile scaling to new domains and tasks.Enhanced simulations and digital twins are intended to provide realistic training scenarios and help bridge the sim-to-real gap.
  • I. Energy Efficiency and Environmental Impact: Modular foundation models, reasoning architectures, and multimodal integration are proposed to improve cross-environment adaptation and causal understanding.The paper identifies plug-and-play adaptation across physical embodiments and environments as a future capability for RFMs and VLAs.
  • I. Energy Efficiency and Environmental Impact: Edge AI, model compression, and specialized hardware are proposed to reduce deployment constraints for real-time inference on mobile robots and embedded systems.Fine-grained control additionally depends on haptic feedback, advanced actuators, and learned dexterity models.
  • I. Energy Efficiency and Environmental Impact: Robust safety protocols and external auditing tools are presented as necessary for monitoring and validating increasingly autonomous systems in critical applications.The proposed frameworks specifically address reasoning failures and control-precision limitations, while continual interactive learning supports improvement through real-world feedback and human interaction.
  • I. Energy Efficiency and Environmental Impact: GPAI combines large-scale generative models with robotic platforms to support adaptive closed-loop behavior that generalizes across tasks and conditions.Multimodal sensing and real-time feedback allow systems to continually refine control policies.
  • I. Energy Efficiency and Environmental Impact: End-to-end training, reinforcement learning, hierarchical planning, and physics-aware simulation support scalable operation in dynamic, high-dimensional settings.Differentiable simulations and physics-aware training target physically plausible predicted behaviors.
Loading 2609.18111v1…