Source-linked AI summary
From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms
Jiangning Zhang, Haojun Chen, Yong Liu
TL;DR
Smart-glasses research is fragmented across devices, tasks, and benchmarks, while reliable deployment requires a complete perception-state-interaction-action loop under resource, privacy, and social constraints. The survey unifies hardware-grounded formalization, capability and application frameworks, and deployment-oriented evaluation, concluding that trustworthy first-person intelligence depends on maintaining useful links among observations, context, intent, and downstream action.
Problem
Smart-glasses research remains fragmented, and isolated model accuracy does not establish a reliable end-to-end loop under wearable operating constraints.
Method
The survey unifies first-person data flow, eight hardware capability axes, seven foundational capabilities, L0-L5 evidential levels, nine application scenes, and claim-conditioned deployment evaluation.
Results
The survey presents a system-level framework linking device profiles, capabilities, application scenarios, deployment design, and claim-conditioned evaluation through a common perception-state-interaction-action loop.
Takeaways & Limitations
Trustworthy deployment requires connecting ongoing first-person observations with spatial and temporal context, user intent, and downstream digital or physical actions.
Takeaways & Limitations
First-person audiovisual recordings typically omit force, tactile, joint-state, and robot-side proprioceptive signals needed for closed-loop control.
Abstract
from arXiv · showhide
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.
1 Introduction
The survey frames smart glasses as first-person intelligence platforms and organizes fragmented evidence into a hardware-grounded, capability-based, application-centered, and deployment-oriented analytical loop.
- Motivation: Smart glasses are analyzed as systems that share the wearer’s viewpoint through integrated sensing, feedback, and interaction rather than as devices users consult periodically.Their everyday positioning links perception and action while avoiding some limitations associated with smartphones and immersive headsets.
- Research gap: The survey addresses fragmentation across augmented reality, egocentric vision, multimodal models, wearable HCI, privacy, and embodied intelligence research.Existing work often relies on offline datasets or abstracts away hardware, runtime, connectivity, bandwidth, thermal, privacy, and social constraints.
- Scope: Its scope is defined by first-person data flow and verifiable system capabilities, distinguishing demonstrated behavior from nominal product features.The survey synthesizes academic, benchmark, prototype, platform, and commercial evidence while using neighboring device classes only as boundary references.
- Contributions: The survey formalizes a first-person observation-to-action loop, constrained utility objective, eight hardware capability axes, and route-aware product profiles.These elements make latency, energy, thermal, privacy, and social costs explicit and convert marketing categories into evidence-bearing experimental substrates.
- Contributions: It organizes seven interdependent capabilities into an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling.Levels are task- and evidence-conditioned, with L5 marking a boundary crossing into embodiment rather than a simple extension of wearer-facing capabilities.
- Contributions: Across nine application scenes, the survey links capability loops with datasets, systems, products, stakeholders, failure consequences, missing evidence, and a deployment-evaluation blueprint.The blueprint covers nine coupled design dimensions, claim-conditioned evaluation, and an evidence ladder from documentation and laboratory measurement through longitudinal deployment and audit.
2 Background
The survey traces smart glasses from early first-person capture and task assistance through shared egocentric data infrastructure and diversified product routes. It formalizes glasses as constrained closed-loop systems and compares hardware using eight capability axes, while showing that routes support different evidence-conditioned claims.
- Evolution: Early smart glasses targeted first-person capture, notifications, remote expertise, procedural guidance, and low-friction information access, but lacked integrated end-to-end loops.The period’s central limitation included missing real-time understanding, sustained feedback, privacy handling, developer ecosystems, and reproducible evaluation.
- Evolution: Egocentric datasets, research platforms, and multimodal models established infrastructure for long-form representation, procedural activity understanding, geometric perception, and interactive assistance.Examples include Ego4D, EPIC-KITCHENS-100, HOI4D, and HoloAssist.
- Product landscape: Recent products diverge into camera/audio-first, camera-and-display, camera-free HUD, and developer or research-oriented routes with different sensing, feedback, and openness profiles.This diversification distributes constrained design resources differently across device capabilities.
- Formalism: The survey defines smart glasses through first-person data flow: aligned observations and intent are transformed into feedback, evolving state, optional actions, and deployment-constrained operation.Inputs include visual, audio, motion, localization, spatial, interaction, and device-state signals; outputs include audio, HUD, AR, or companion-device feedback.
- Formalism: The constrained objective maximizes real-world task utility while explicitly limiting latency, energy, thermal load, privacy risk, and related deployment costs.The formulation treats evidence, authority, operating conditions, versions, and task risk as part of deployment-oriented evaluation.
- Hardware framework: Eight capability axes provide a common coordinate system linking hardware and platform interfaces to application requirements and standardized evaluation.The survey cautions that progression is neither monotonic in sensor count nor proportional to display complexity; camera/audio-first products provide near-term L2 evidence, while stronger claims require additional verification.
- Scope boundaries: Embodied data-interface systems offer methodological references for collecting human demonstrations and multimodal behavior data, but remain outside the survey’s smart-glasses scope.These systems support embodied-intelligence datasets, robot imitation learning, and in-the-wild human-data management.
3 Foundational Capabilities for Smart Glasses
The survey organizes first-person intelligence as a dependency structure rather than a linear pipeline, connecting perception to multimodal context, spatial state, memory, action, and embodied interfaces. It emphasizes temporally reliable, provenance-aware, interaction-sensitive capability evaluation.
- Framework: The capability structure links first-person perception, multimodal context, spatial state, personal memory, and situated action, with embodied interfaces extending it toward cross-embodiment transfer.Deployment constraints cut across all capabilities, while L0-L4 describe wearer-facing progression and L5 is orthogonal.
- First-person perception: First-person perception supplies evidence aligned with wearer behavior, task context, and ongoing interaction rather than merely recognizing isolated entities.Egocentric framing, continuous motion, hand occlusion, and interactive procedures create distribution and evaluation challenges.
- Evaluation: Evaluation must extend beyond isolated recognition toward wearable reasoning, streaming interaction understanding, real-time text recognition, and contextual assistance.The cited benchmarks collectively test perception under device and interaction conditions rather than only frame-level accuracy.
- Multimodal context: Multimodal context must align video, speech, sound, motion, gaze, location, device state, history, and feedback while preserving modality-specific uncertainty.Wearable question answering, speaker-aware fusion, and long-video retrieval motivate explicit contextual state rather than independent evidence collections.
- Temporal state: Temporal state should preserve event order, validity, task phase, unresolved dependencies, and evidence responsible for each update as situations evolve.Streaming interaction understanding and query-conditioned retrieval support selective access to evidence distributed across long observation streams.
- Interaction: Bidirectional interaction makes feedback part of reasoning because acoustic processing and everyday conversations jointly involve environmental awareness, intelligibility, interruption, misunderstanding, and repair.Output channels can therefore affect both what the wearer hears and how the system interprets interaction.
- Provenance: Contextual state should distinguish observations, inferences, retrieved memories, and external-tool outputs while handling conflicting evidence and attribution across views, speakers, and modalities.Cross-view memory, H2HMem, OpenEQA, and SAW-Bench motivate provenance-aware grounding.
3.4 Auditable Long-Term Personal Memory
Auditable long-term personal memory must preserve provenance, temporal and spatial validity, user control, and authorization across the full memory lifecycle. Its evaluation therefore extends beyond retrieval accuracy to correction, revocation, deployability, and action governance.
- Memory organization: Personal memory must distinguish working, episodic, semantic, spatial, user-confirmed, and action-log information with different temporal horizons and authorization policies.The proposed organization treats memory classes as having distinct retention and access requirements.
- Provenance and admission: Memory items should retain source, timestamp, spatial context, confidence, confirmation status, access permissions, and expiration policy.Admission rules should distinguish direct observations, model inferences, external retrieval, and user-confirmed facts.
- Temporal validity: Useful recall must recover when, where, and under what evidence an assertion was established, not only its semantic content.The cited benchmarks target distant-event reasoning, evolving episodic streams, everyday assistance, and long-term identity preservation.
- Correction and revocation: User-controllable memory requires provenance inspection, correction, person- or place-specific deletion, permission changes, and verification that updates affect later behavior.Consent and revocation may depend on social context involving bystanders, not only the wearer.
- Evaluation: Memory evaluation should jointly test relevance, evidential support, validity, user control, and availability within device resource limits.Individualized context can improve visual understanding, while online episodic question answering exposes latency and resource constraints.
- Connection to action: Context and memory provide evidence for action but do not confer execution authority without authorization, outcome monitoring, correction, and recovery.Situated action is defined as a governed observation-state-action loop rather than answer generation alone.
3.7 Cross-Cutting Deployment Constraints
Deployment constraints define the feasibility envelope for every smart-glasses capability, spanning computation, energy, thermal behavior, human factors, latency, privacy, security, service dependence, and reproducibility. The L0-L5 framework makes these boundaries claim- and evidence-dependent rather than product-wide scores.
- Computation and energy: Streaming inference, adaptive sampling, cascaded models, on-device filtering, and edge-cloud scheduling jointly shape observation quality, latency, energy consumption, and sustainable duty cycle.Efficient egocentric perception and visual-inertial odometry illustrate the need to co-design computation with constrained hardware.
- Wearable operating envelope: Battery, thermals, weight, camera placement, display brightness, audio leakage, network dependence, and firmware can substantially change runtime behavior.Assistive recognition performance is materially affected by walking speed, camera placement, and camera type.
- Latency and feedback: A semantically correct response can fail when it arrives too late, is overly detailed, uses inappropriate volume, or obstructs the wearer’s field of view.Real-time response and conversational timing are system-design objectives, not merely model-output properties.
- Privacy and security: Privacy-preserving operation must cover acquisition, filtering, transmission, storage, retrieval, action, and revocation, with controls for consent, access, adversarial robustness, and auditability.Threats include visual attacks, social engineering, and failures of unlearning-based defenses.
- Evidence boundaries: Capability levels are assigned to demonstrated regimes under stated tasks, hardware, operating conditions, and evidence, not inferred from sensor counts, displays, brands, or aggregate product scores.Qualifiers such as L3 potential, L5 data, and partial L5 make incomplete public substantiation explicit.
- L0-L5 framework: L0-L4 progress from capture and reactive perception to contextual assistance, persistent state, and governed action, while L5 separately denotes embodied coupling across the embodiment boundary.Each level has distinct minimum conditions and evaluation targets.
4 Application Scenes for Smart Glasses
The application-scene analysis shifts from isolated system functions to deployable requirements under concrete user activities. It examines how sensing, reasoning, memory, feedback, and action combine in real-world use.
- Application-scene perspective: Application scenes translate combinations of sensing, reasoning, memory, feedback, and action into requirements tied to concrete user activities.The analysis moves beyond isolated functions toward deployment-oriented evaluation.
4.1 Daily Situated Assistance
Daily situated assistance connects the wearer’s current first-person view, personal context, and external services for low-friction support. Moving from perception to memory and execution raises requirements for grounding, timing, permission, recovery, privacy, and repeated-use utility.
- Scene scope: Daily assistance covers reading, translation, conversation, retrieval, object finding, reminders, meetings, travel, retail queries, and creative capture.Camera- and audio-first products provide practical entry points for user-initiated L2 assistance.
- Perceptual access: Immediate assistance must convert a moving wearer-centered audiovisual stream into concise, correctly grounded outputs through available audio or visual channels.Relevant functions include reading signs, translating language, answering nearby-object questions, and supporting conversations without a phone.
- Personal context and memory: Persistent assistance links current observations with personal history, including where objects were last seen, earlier discussions, unfinished tasks, and contextually relevant reminders.EgoLife, personalized visual context, and ContextAgent represent routes toward this capability.
- Agentic execution: Action-capable assistance connects first-person context to web search, communication, bookings, purchases, and other services, but requires explicit permissions and cancellation or rollback.Payments, bookings, data sharing, and other consequential actions should be evaluated separately from low-risk information retrieval.
- Evaluation: End-to-end evaluation should report faithfulness, grounding, tail latency, interruption burden, false activation, correction, privacy leakage, memory provenance, action recovery, and repeated-use utility under varied operating conditions.Conditions include background speech, changing illumination, partial visibility, intermittent connectivity, limited battery, and service updates.
4.2 Accessibility Assistance
Accessibility assistance uses smart glasses for visual, auditory, communicative, mobility, and cognitive support, but needs vary substantially across users and activities. Evaluation therefore must extend beyond average accuracy to independent, effort-aware, personalized outcomes.
- Accessibility applications support textual, environmental, communicative, and mobility-related information access for heterogeneous users.Relevant needs include low-vision and blindness assistance, hearing access, cognitive cueing, and mobility guidance.
- Visual access and environmental understanding: Camera-equipped glasses can provide scene description, text reading, object and person recognition, product identification, and visual-detail enhancement.VizWiz, EgoBlind, and OCR-Wearable address questions, continuous guidance, and the effects of walking speed and camera configuration.
- Communication access and auditory awareness: Lightweight visual-cueing products support live captions, speech translation, speaker identification, and acoustic-event alerts in everyday communication.Camera/audio systems can additionally use visual context to identify speakers or disambiguate references.
- Mobility, hazard awareness, and cognitive cueing: Navigation and cognitive support require longer, safety-sensitive loops that connect hazard understanding, active sensing, reminders, and mobility guidance.NavCog, LidSonic, and UrbanRiskVQA provide references for navigation, active sensing, and urban-risk understanding.
- Personalization, independent outcomes, and responsible deployment: Average accuracy cannot represent assistive value across populations and activities; evaluation should measure independence, effort, demand, recovery, near misses, and personalization cost.The proposed outcomes include independent task completion, time and effort, cognitive load, visual and auditory demand, error recovery, near-miss events, and personalization cost.
4.3 Industrial Workflow Support
Industrial workflow support embeds smart glasses in SOP-constrained processes where permissions, quality, safety, and audit requirements shape assistance. Useful systems must guide procedures, document operations, support recovery, and withstand real workplace conditions without enabling unnecessary surveillance.
- Industrial applications cover assembly, maintenance, inspection, warehouse work, laboratory operations, and field service within SOP- and role-constrained processes.These environments impose quality targets, safety rules, station- and role-specific permissions, and audit responsibilities.
- SOP-grounded execution and procedural guidance: Procedural assistance must identify the current step, relevant objects or tools, completed prerequisites, and next permissible action.Assembly101, IKEA-ASM, COIN, CrossTask, LabOS, and HoloAssist provide complementary resources for procedural and situated assistance.
- Inspection, identification, and operational documentation: Inspection and field workflows add identification, equipment-state grounding, quality checks, evidence capture, incident review, and searchable audit logs.Data minimization and role-based access are necessary to prevent continuous workplace capture from becoming unnecessary worker surveillance.
- Remote expertise, proactive intervention, and error recovery: First-person streaming enables remote experts to view task context, annotate regions, and guide recovery while leaving workers’ hands free.HoloAssist, Pro2Assist, Plan-Watch-Recover, Streaming Interventions, and IPIBench address interactive, proactive, planning, monitoring, and recovery functions.
- Field robustness, governance, and organizational outcomes: Industrial validation must cover noise, gloves, dust or oil, temperature, occlusion, protective equipment, intermittent connectivity, and outdated or conflicting information.These conditions define the operating envelope in which workers actually use the glasses.
4.4 Healthcare and Caregiving
Healthcare and caregiving span documentation, education, rehabilitation, reminders, elder care, and possible decision support, with consequences increasing as systems approach clinical recommendation. Deployment therefore requires workflow validation, professional oversight, and jointly managed privacy, liability, hygiene, security, and audit constraints.
- Healthcare and caregiving applications range from clinical capture and education to rehabilitation, reminders, elder care, and potential decision support.The same perception or cueing mechanism can have different consequences depending on whether it records, reminds, or influences diagnosis or treatment.
- Clinical workflow capture and documentation: Hands-free first-person capture can document procedures, preserve clinicians’ views, support review, and reduce manual note-taking.Cholec80/EndoNet and JIGSAWS provide references for surgical-phase recognition and skill or action analysis, but do not alone validate wearable clinical deployment.
- Professional education and procedure-aware assistance: Medical education and supervised training are more strongly supported than autonomous clinical decision making.AR medical-education reviews, HoloAssist, and Ego-Exo4D motivate demonstration replay, viewpoint sharing, step-aware prompts, and post-hoc review.
- Rehabilitation, elder care, and home support: Home and long-term care can use smart glasses for reminders, activity coaching, object or medication finding, caregiver communication, and accessible perception.Envision Glasses, restorative augmented-reality narratives, EgoLife, and PVCL provide relevant product and research routes.
- Clinical evidence, PHI governance, and responsibility boundaries: Clinical claims require professionally annotated workflow logs, representative-condition validation, PHI governance, consent, liability, hygiene, security, and audit trails.Professional confirmation and override should be reported as system outcomes because errors may propagate into regulated decisions, records, or care responsibilities.
4.5 Education and Skills Training
Education and skills training use smart glasses to support durable learning through demonstration capture, step-aware coaching, reflection, and personalization. Valid evaluation must distinguish immediate task completion from retention, transfer, independent performance, agency, and governed data use.
- Educational applications target durable knowledge and skill acquisition across laboratories, cooking, repair, sports, music, language, creative practice, and vocational training.The loop can capture demonstrations, decompose steps, estimate learner state, select feedback, and support later reflection.
- Demonstration capture and skill decomposition: First-person and multiview recordings expose expert attention, object manipulation, and procedural progression for demonstration capture and skill decomposition.Ego-Exo4D, HoloAssist, Ego-1K, COIN, CrossTask, Assembly101, and IKEA-ASM provide related resources.
- In-situ coaching and error-aware practice: In-situ coaching can prompt, detect missed or incorrect steps, demonstrate corrections, or defer to a teacher while tracking task progress and intervention timing.Pro2Assist, Plan-Watch-Recover, Streaming Interventions, and IPIBench motivate task-aware assistance.
- In-situ coaching and error-aware practice: Immediate correction may improve short-term task success, whereas delayed hints, questions, or fading support may better promote recall and problem solving.Evaluation should report intervention precision, recovery, cognitive load, prompt dependence, and ability to continue afterward.
- Reflection, personalization, and inclusive learning: Reflection tools should make interpretations inspectable and allow learners or teachers to correct task state, annotate mistakes, and select retained information.Personalized support must include calibration cost and risks of inappropriate adaptation in evaluation.
- Retention, transfer, agency, and data governance: Existing resources mainly capture demonstrations or in-task assistance, leaving longitudinal evidence on retention, transfer, independent performance, agency, and governance incomplete.Studies should distinguish immediate completion from delayed transfer and performance without the glasses, while addressing prompting and access rights.
4.6 Mobility and Transportation Safety
Mobility and transportation safety require smart glasses to perceive hazards and communicate guidance within limited decision windows without masking environmental awareness. Evaluation must measure both system performance and user response across realistic motion, weather, illumination, and traffic conditions.
- Scope: Mobility scenes combine wayfinding, hazard detection, traffic-intention understanding, outdoor sports feedback, and emergency alerts under continuous motion.Relevant tasks include route progress, turn selection, hazard detection, pedestrian intention, and performance feedback.
- Wayfinding and route-level guidance: Route guidance must synchronize localization, route progress, turn selection, destination awareness, and deviation recovery with intelligible feedback.NavCog and UrbanRiskVQA provide reference resources, but wearable guidance must remain aligned with wearer position and movement.
- Hazard detection and traffic-intention understanding: Safety assistance should report missed hazards and false alarms alongside reaction time, inappropriate responses, alert habituation, and near-miss events.Third-person road datasets provide useful references, but head-mounted sensing creates a non-direct transfer problem and asymmetric error costs.
- Sports mobility and outdoor multimodal interaction: Outdoor operation depends jointly on camera placement, ingress protection, battery endurance, microphone robustness, display legibility, and audio masking.Evaluation should include motion degradation, wind and traffic noise, weather, nighttime conditions, battery depletion, and awareness of surrounding hazards.
- Safety-critical evaluation and operating boundaries: A unified safety protocol remains absent, so evaluation should separately stratify warning latency, reaction time, misses, false alarms, distraction, masking, near misses, and recovery behavior.Recommended stratification factors include illumination, weather, route complexity, traffic density, and user mobility characteristics.
4.7 Social Interaction and Collaboration
Social interaction and collaboration use smart glasses to mediate dialogue, share viewpoints, preserve task state, and coordinate multiple stakeholders. Success depends not only on technical correctness but also on permissions, participation, privacy, and longitudinal acceptance.
- Scope: Collaborative glasses support captions, translation, telepresence, decision memory, and role handoffs across wearers, experts, peers, teachers, patients, customers, and bystanders.Unlike individual assistance, these scenarios depend on what multiple stakeholders understand and what information each may access.
- Meeting and dialogue mediation: Meeting assistance can provide captions, translation, speaker-turn cues, query-focused summaries, and reminders of unresolved points.AMI, QMSum, and ConversationBreakdowns provide complementary meeting, summarization, and everyday wearable-interaction resources.
- Audio-visual social perception and inclusive participation: Inclusive collaboration requires identifying speakers, tracking directed attention, and sharing information across participants with different perceptual abilities.AVA-ActiveSpeaker, MixedVision, and AR-DeafEducation support complementary audio-visual and accessibility-oriented directions.
- Remote collaboration, shared viewpoints, and expert handoff: First-person streaming enables remote experts to observe work contexts, indicate relevant regions, and guide wearers through procedures.HoloAssist, Pro2Assist, Plan-Watch-Recover, and enterprise products represent research and implementation routes for procedural assistance.
- Shared memory and role-aware coordination: Stateful collaboration requires distinguishing what the wearer observed from what another participant said, inferred, or committed.H2HMem, EgoExoMem, EgoMemReason, and PVCL address multimodal, cross-view, long-horizon, and personalized memory routes.
- Consent, bystander privacy, and longitudinal social acceptance: Camera-glasses deployment must address bystander consent, privacy-utility trade-offs, group-based controls, and longitudinal social acceptance.The evidence base includes CameraGlassesPrivacy, BystanderPrivacy, MindTheGap, PrivacyUtility, VisGuardian, and social-acceptability studies.
4.8 Spatial Intelligence
Spatial intelligence extends smart glasses from momentary perception to persistent representations of objects, places, paths, and actionable regions. Reliable operation requires a chain from calibrated mapping and reference to maintained world state and governed action.
- Scope: Spatial intelligence covers indoor wayfinding, object finding, spatial reminders, digital twins, guidance, shared annotations, and handoff to IoT devices or robots.These tasks operationalize persistent spatial state across everyday and shared environments.
- Spatial reconstruction and semantic mapping: Spatial reconstruction combines first-person cameras, inertial measurements, and calibration to estimate pose, reconstruct geometry, and attach semantic labels.Wearable maps also require evaluation of drift, relocalization, dynamic objects, partial coverage, sensing duty cycle, timestamps, provenance, and uncertainty.
- Navigation and spatial question answering: Localized glasses can guide routes and answer questions about object locations, accessible entrances, and destinations.Habitat, R2R, REVERIE, EmbodiedQA, and OpenEQA provide navigation, reference, instruction-following, and embodied-question-answering resources.
- Deictic reference, gaze, and near-body relations: Deictic spatial interaction depends on linking expressions such as “this” or “over there” with gaze, pointing, proximity, and hand-centered relations.PointingMLLM/EgoPoint-Bench, EgoProx, and EGTEA Gaze+ cover complementary referential, proximity, gaze, and action reasoning.
- Cross-session world state and digital twins: Persistent spatial services must detect stale maps, relocalize across sessions, timestamp state, expose confidence, support correction and rollback, and enforce access rules.Objects move, spaces change, permissions differ, and old maps can remain locally accurate while becoming operationally wrong.
- Spatial action interfaces and end-to-end validation: Spatial action interfaces connect maintained state to world-locked annotations, device interfaces, robot targets, or IoT services.An L4 spatial-action claim must specify the target and coordinate relationship rather than stopping at spatial understanding.
4.9 Embodied Intelligence
Embodied intelligence treats smart glasses as interfaces for capturing human experience, structuring demonstrations, and transferring task knowledge to robots. The central boundary is that visual egocentric data alone may not establish contact, force, executable equivalence, or safe downstream outcomes.
- Scope: Smart glasses can capture demonstrations, narratives, failures, recoveries, gaze, body and hand motion, hand-object relations, and spatial pose while supporting robot supervision.This creates a longer pipeline than ordinary wearer assistance, extending from sensing and feedback to transfer and physical execution.
- First-person demonstrations and multiview data infrastructure: Useful embodied-data pipelines synchronize devices, calibrate coordinate frames, segment activities, link actions to object-state changes and outcomes, and measure usable data per hour.They must also preserve consent coverage, de-identification, workspace privacy, and ownership of demonstrated skills.
- Active vision and embodied state estimation: Head motion can provide active information-seeking signals and help identify viewpoints useful for manipulation.ActiveGlasses, EgoMI, ActiveMimic, EPIC, and LEVIO span active perception, manipulation, egocentric pretraining, and resource-constrained estimation.
- Contact, affordance, and dexterous skill recovery: Robot learning must infer active objects, contact locations, object-state changes, affordances, and human subgoals rather than merely recognize visible actions.EgoDex, EgoScale, EgoEngine, UniDex, EgoTactile, and ForceBand address complementary dexterous and force-related recovery problems.
- Contact, affordance, and dexterous skill recovery: Ordinary smart glasses do not directly observe force, tactile feedback, joint torque, or full proprioception, so visual contact estimates require uncertainty and downstream physical validation.Retargeting must account for human–robot differences in hands, reachability, kinematics, and safety constraints.
- Human-to-robot representation and policy transfer: Credible human-to-robot transfer must align viewpoints and task state, convert human motion into robot action space, and validate outcomes on the robot side.The literature spans imitation, zero-shot transfer, human-video representations, VLA policies, robot datasets, and policy or affordance-grounded control.
- Robot-side validation, safety, and governance: An L5 system claim is justified only when the complete chain from human capture to safe robot outcome is substantiated.Responsibility, consent, de-identification, data-use boundaries, and skill ownership remain part of evaluation for demonstrations collected in private or proprietary settings.
4.10 Other Potential Application Scenes
Smart glasses extend beyond the nine major application scenes into emerging uses such as tourism, retail, sports, creative production, and entertainment, but evidence remains distributed across products, prototypes, and adjacent research traditions. Across scenes, deployment requirements vary with observability, persistence, action consequences, stakeholders, and available validation evidence.
- Other Potential Application Scenes: Emerging application areas include tourism, cultural-heritage interpretation, museum guidance, location-aware retail, sports coaching, creative production, and shared augmented-reality entertainment.These applications currently draw evidence from products, prototypes, and neighboring research traditions.
- Other Potential Application Scenes: Hardware profiles constrain what smart glasses can observe, compute, remember, and communicate to the wearer.The constraint applies across application scenes and depends on the device’s available sensing, processing, memory, and feedback pathways.
- Other Potential Application Scenes: State persistence, action consequences, stakeholder structure, and evidence availability determine how deployment requirements and validation thresholds vary by scene.No foundational capability has a universal pass criterion; requirements must reflect task risk and deployment conditions.
5 Design Framework and Evaluation
The design framework treats smart glasses as closed-loop embodied assistance systems whose hardware, runtime, state, feedback, action, reliability, governance, and reproducibility must be evaluated together. Its claim-conditioned protocol and iterative evidence ladder tie deployment conclusions to specific tasks, configurations, risks, participants, environments, and evidence types.
- Overall Design Framework: The closed-loop embodied assistance process, rather than an individual model, is the primary design object for deployable smart glasses.The system must repeatedly decide what to sense, which observations support reasoning, where computation occurs, what state persists, and how action is authorized.
- Hardware Profile: Hardware determines the physical envelope for observation, multimodal fusion, feedback, computation, communication, power, thermals, and sustained wear.Sensor geometry, calibration, timestamp integrity, comfort, social acceptability, and endurance all affect the usable platform.
- Hardware Profile: There is no universally optimal form factor: navigation, memory assistance, industrial guidance, and embodied-data platforms prioritize different hardware capabilities.Hardware claims should report configuration, alignment, calibration, fit, power, thermal behavior, recording signals, and longitudinal adherence, supported by measurement and field evidence.
- Runtime Design: Runtime topologies distribute sensing, computation, state, and services across glasses, phones or pucks, clouds, local servers, and on-device execution.Placement changes data exposure, latency, availability, energy use, state consistency, failure responsibility, and degraded-connectivity behavior.
- Persistent State and Memory Lifecycle: Persistent-state writes require task-dependent authorization or user confirmation, while stored items should remain inspectable, correctable, scope-limited, or rejectable.The policy is especially important for identities, preferences, commitments, health-related observations, and descriptions of other people.
- Feedback and Intervention: Feedback quality depends on correctness together with modality, timing, duration, referent, and initiative, because poorly grounded interventions can burden users or induce incorrect action.A nominally accurate prediction can still produce harmful interaction when delivered at the wrong time or without adequate grounding.
- Structured Evaluation Protocol: Standardized evaluation asks whether a particular capability claim is supported under stated task, hardware, runtime, participant, and risk conditions rather than ranking heterogeneous products universally.The protocol treats the nine design dimensions as coupled evaluation objects within a shared claim and reproducibility context.
- Structured Evaluation Protocol: The atomic evaluation unit conditions each capability claim on task, hardware route, sensing and output, L0-L5 level, risk, participants, environment, evidence type, and confidence.N/A means unavailable evidence rather than zero capability, and benchmark or prototype results should not be generalized to longitudinal field use without corresponding evidence.
6 Conclusion and Future Prospects
Smart glasses remain difficult to deploy as reliable embodied-intelligence platforms because they must sustain an end-to-end perception-state-interaction-action loop under hardware, data, personalization, timing, interoperability, and evaluation constraints. The survey proposes reproducible, auditable, and adaptive system designs to address these boundaries.
- System-level challenge: Reliable smart-glasses intelligence requires continuous observation, temporally valid state, appropriate intervention, communication, and recoverable action rather than isolated model accuracy.The central challenge is sustaining the complete loop under wearable operating conditions.
- System-level challenge: Weight, battery, thermal, optical, sensing, synchronization, and connectivity constraints jointly determine whether closed-loop operation remains useful beyond short demonstrations.These constraints can propagate through sensing and inference quality, state freshness, and interaction reliability.
- Data and personalization: Longitudinal first-person data remain constrained by bystander privacy, sensitive locations, sparse critical events, annotation cost, device-specific views, calibration drift, and hardware or firmware changes.These limitations affect relocalization, memory, preference learning, retrieval, and long-term world-state validity.
- Data and personalization: Personalization must account for differences in ability, language, motion, culture, workload, and interruption tolerance while preventing overfitting, sensitive-trait inference, and group-specific regressions.The survey calls for explicit correction and preference control alongside group-stratified evaluation.
- Interaction and action: Proactive assistance must jointly optimize intervention timing, modality, duration, specificity, confirmation, override, and rollback against task risk and individual preference.Evaluation should include missed-intervention cost, false-alarm burden, attention disruption, recovery time, and repeated-error rate.
- Future roadmap: Evaluation should condition claims on device, sensors, software versions, operating conditions, task scope, temporal horizon, action authority, and risk level, combining controlled replay with real-device field testing.The survey also calls for latency and energy distributions, recovery, uncertainty, privacy, and failure reporting.
- Future roadmap: The survey recommends versioned device profiles, privacy-aware longitudinal data engines, auditable memory, adaptive interaction, interoperable action ecosystems, and robot-validated transfer.These directions are intended to make system behavior inspectable, reproducible, correctable, and deployable.
- Interaction and action: Interoperable systems need shared schemas, capability negotiation, and transactional execution that separates proposal, confirmation, execution, verification, and rollback.The display can serve as a confirmation and provenance surface, while spatial and tool results should be reconciled before later actions.