Source-linked AI summary
Agentic UE-CoMIMO for 6G Terminals: From Virtual Antenna Augmentation to AI-Native Virtualization
Chao-Kai Wen, Yen-Cheng Chan, Lung-Sheng Tsai, Pei-Kai Liao, Geoffrey Ye Li
TL;DR
Individual UEs face antenna-scaling limits, while extending UE-CoMIMO across communication, sensing, computing, and semantic exchange requires control that can interpret intent and replan. The paper introduces a hierarchical Agentic UE-CoMIMO system coordinating devices, mechanisms, computation, and tokens. Across live-streaming and blind-spot-sensing scenarios, anticipation sustains quality longer and preserves warnings through device outages.
Problem
UE terminals have limited antennas, and coordinating multi-device communication, sensing, computing, and information exchange requires control beyond fixed objectives and actions.
Method
Agentic UE-CoMIMO uses device micro-agents, a smartphone or CPE hub, and edge/network agents to coordinate participation, modes, computation, semantic tokens, and topology.
Results
Agentic control sustains high-quality streaming longer and maintains blind-spot warnings through device outages, while using 8.8 Mbps average encoded bitrate versus 18.6 Mbps for the myopic controller.
Takeaways & Limitations
Prediction, intent interpretation, and feedback-driven replanning improve longer-session resource sustainability and service continuity beyond cooperation and instantaneous adaptivity alone.
Abstract
from arXiv · showhide
End-user-centric collaborative MIMO (UE-CoMIMO) lets nearby devices form a virtual multi-antenna terminal to overcome the antenna limitations of individual user equipment. Extending such cooperation to communication, sensing, computing, and task-relevant information exchange requires a control layer that can interpret user intent, select cooperation mechanisms, and replan as conditions change. This article introduces Agentic UE-CoMIMO, in which device micro-agents, a smartphone or CPE hub agent, and edge/network agents coordinate device participation, relay modes, traffic splitting and duplication, compute placement, semantic-token exchange, and topology reconfiguration. Two system-level scenario studies on creator-centric live streaming and wearable-collaborative blind-spot sensing compare the proposed controller with capability-matched adaptive baselines. The results show that, by anticipating changes and preparing cooperation and fallback actions in advance, agentic control sustains high-quality streaming for longer and maintains blind-spot warnings through device outages. We also discuss the associated standardization, interoperability, trust, and validation challenges.
I. INTRODUCTION
UE-side antenna scaling is constrained by terminal form factor, power, spacing, and thermal budgets, motivating UE-CoMIMO’s virtual antenna cooperation. Agentic UE-CoMIMO adds user-side control that interprets intent, coordinates heterogeneous mechanisms, and replans as tasks and device states change.
- UE terminals cannot support additional spatial streams when form factor, antenna spacing, power, and thermal limits constrain their antenna count.
- UE-CoMIMO recruits nearby collaborative UEs and local infrastructure to form a virtual antenna system, with demonstrated gains in effective channel dimensionality and throughput.Frequency-translation relays and relay-assisted carrier aggregation provide physical cooperation mechanisms.
- Agentic UE-CoMIMO interprets intent, decomposes tasks, selects cooperation mechanisms, maintains task state, and replans from observed outcomes.The control layer supports communication, sensing, and computing objectives across changing device and channel conditions.
- Agentic control differs from conventional adaptation because the objective, participating devices, and required mechanisms can change during task execution.Conventional optimization or learned policies may still operate inside a mechanism selected by the agentic layer.
- Unlike primarily network- or AP-side control, Agentic UE-CoMIMO places control on the user side and jointly manages device roles, physical cooperation, sensing, computing, and semantic exchange.The coordination occurs under antenna, battery, thermal, trust, and radioresource constraints.
II. AGENTIC UE-COMIMO: ARCHITECTURE FOR AI-NATIVE VIRTUALIZED TERMINALS
Agentic UE-CoMIMO extends physical-layer virtual antenna aggregation into an AI-native virtualized terminal spanning communication, sensing, computing, and semantic exchange. Its control objective includes task quality and device-state constraints alongside spectral efficiency and spatial rank.
- The architecture extends UE-CoMIMO from physical-layer antenna cooperation to coordinated communication, sensing, computation, and semantic exchange.
- UE-CoMIMO treats nearby user-side devices as a cooperative virtual antenna system, including devices that may belong to different trust domains.Personal clusters are easier to authorize, while CPEs, access points, vehicles, and roadside nodes may offer better assistance but raise authorization and incentive issues.
- Physical cooperation can provide diversity, higher-rank effective MIMO channels, or extended virtual arrays for localization.L1 frequency-translation relays suit latency-critical, high-rate services; L2/L3 relays support packet processing and cross-RAT routing for less time-sensitive services.
- Agentic UE-CoMIMO adds agentic control, distributed computing, semantic-token exchange, and integrated communication, sensing, and perception while retaining physical-layer mechanisms.The control objective includes quality of experience, sensing confidence, energy consumption, and thermal state in addition to spectral efficiency and rank.
B. Hierarchical Agentic AI
The hierarchical architecture distributes intelligence across constrained devices, a real-time smartphone or CPE hub, and slower but better-resourced edge/network support. This separation enables local awareness and coordination while keeping decisions within network-defined limits.
- Lightweight devices host constrained micro-agents, while smartphones or CPEs serve as hub agents and edge/network agents provide broader resources and visibility.
- Micro-agents report link quality, battery, temperature, traffic deadlines, and locally derived features or tokens, requesting fallback actions when conditions become critical.
- The smartphone or CPE hub selects participating devices and roles, relay or traffic modes, compute placement, and semantic-token budgets.It acts as the virtualized terminal’s control plane without processing every payload.
3) Edge/Network Agent:
Agentic UE-CoMIMO uses edge/network support for policy, maps, coordination, and large-model capabilities while the local hub makes fast decisions through a closed-loop agentic controller. The controller can reorganize the control problem as intent and operating conditions evolve.
- The edge/network agent supplies radio and sensing maps, inter-user coordination, large-model inference, and a policy envelope defining resource, parameter, and security limits.
- The control loop observes context, interprets intent, selects topology and mechanisms, exchanges semantic information, monitors task quality, and revises decisions.
- Agentic control can interpret intent, decompose tasks, combine heterogeneous mechanisms, maintain state, monitor outcomes, and revise decisions when objectives or conditions change.
- Evaluation emphasizes prediction, intent interpretation, and feedback-driven replanning as distinct capabilities.These capabilities use expected system evolution, service-specific tradeoffs, and device-state or downstream-quality feedback.
- The evaluations use interpretable rules and state machines with capability-matched baselines lacking forward planning, intent interpretation, and feedback-driven replanning.
A. Intent-Aware Multi-Device Coordination
The hub agent selects cooperating devices, assigns roles, and configures relay, traffic, computation, and token resources according to the current service objective and user intent.
- At each decision instant, the hub selects a cooperating subset, assigns device roles, and sets relay, traffic, compute, and token configurations.
- Configuration choices balance communication, sensing, and perception quality against energy use, coordination overhead, and thermal stress.
- The same device cluster can adopt different configurations for throughput-oriented upload, latency-critical interaction, or safety-critical sensing.
- Because intent may change during a session, the control objective and cooperating configuration can change accordingly.
B. Semantic-Token Aggregation
Agentic UE-CoMIMO uses task-aware semantic tokens as an interoperable interface for heterogeneous devices, while allowing topology and information exchange to adapt as task quality changes.
- Semantic tokens replace costly raw observations with compact, task-aware payloads such as channel features, sensing confidence, or blockage indicators.
- Each token combines structured features or an embedding with metadata describing task, source, uncertainty, freshness, trust, and resolution.
- Token generation can use thresholding or CFAR-like processing on constrained devices and learned encoders on more capable devices.
- The device cluster is reconfigured across live streaming, sensing-token exchange, and safety-warning use cases.
- Task-quality feedback can trigger recruitment, higher-resolution tokens, critical-packet duplication, or computation migration when confidence, latency, or thermal conditions change.
IV. TWO REPRESENTATIVE USE CASES
The paper presents two representative use cases: creator-centric live streaming assisted by CPE and wearable-collaborative blind-spot sensing and warning.
- The two use cases combine creator-centric live streaming with CPE assistance and wearable-collaborative blind-spot sensing and warning.
A. Creator-Centric Live Streaming with CPE Assistance
The live-streaming study evaluates agentic coordination as a creator moves from outdoor coverage into an indoor RF dead zone. Cooperation and prediction sustain target-quality streaming while reducing resource use relative to a capability-matched myopic controller.
- The scenario models first-person streaming from AI glasses during an outdoor-to-indoor transition, where outdoor blockage and indoor penetration loss or congestion create different bottlenecks.
- The hub can recruit a smartphone or puck, split or duplicate traffic, move transmission or computation, and use CPE as relay, receiver, or computing host.
- The control objective covers visual quality, tail latency, energy consumption, and stream continuity across the transition.
- The evaluation compares fixed and reactive cellular baselines, a no-predictive-governor ablation, and a capability-matched myopic controller against the proposed policy.
- All policies with full cooperation mechanisms continue streaming through the RF dead zone, while cellular-only baselines stall or fail to recover.
- 18.6 Mbps versus 8.8 Mbps average encoded bitrate: the myopic controller achieves nearly the proposed policy’s instantaneous QoE but uses more resources.
B. Wearable-Collaborative Blind-Spot Sensing and Warning
The wearable sensing study uses bistatic smartwatch–CPE cooperation and semantic-token exchange to preserve rear-hazard warnings as blockage and CPE failure occur. Agentic recovery outperforms static and reactive alternatives while reducing reliance on the always-on smartwatch.
- System design: A smartphone hub coordinates a rear-facing bistatic sensing path in which the smartwatch or scooter-mounted CPE receives echoes while streaming remains active.The hub can switch to safety priority, limit video enhancement, reserve sensing resources, and duplicate warning packets over both uplinks.
- System design: Semantic-token aggregation lets assisting devices send compact information that the hub fuses into warnings containing confidence, direction, and urgency.When confidence decreases, the hub can recruit another device or reassign transmitter and receiver roles.
- Evaluation: The simulation tests smartwatch blockage followed by temporary CPE unavailability against no cooperation, static delegation, reactive control, and agentic control.The smartwatch is the default receiver and the scooter-mounted CPE is the auxiliary receiver; the smartwatch sensing chain can otherwise be disabled for energy efficiency.
- Evaluation: The agentic hub immediately returns to smartwatch-assisted sensing after CPE failure, whereas static delegation loses warnings and reactive control produces a gap exceeding one second.During blockage, the cooperating policies use the CPE, whose detection probability remains near 0.84, and warn within approximately 0.1 s after target entry.
- Evaluation: 98.7% warning reliability is achieved by the agentic policy, versus 82.9% for reactive control, 65.8% for static delegation, and 56.6% for no cooperation.The agentic policy also achieves the highest joint QoE while consuming less wearable energy than the always-on-watch baseline.
- Broader scope: The same coordination problem extends to field service, construction safety, and immersive event broadcasting under device and network constraints.
V. STANDARDIZATION AND OPEN RESEARCH CHALLENGES
Agentic UE-CoMIMO requires standardization and governance beyond physical-layer cooperation, especially for bounded local control, interoperable tokens, trust, privacy, and validation under realistic impairments.
- Standardization: Table I separates existing mechanisms, implementable extensions, and genuinely new standardization requirements across Agentic UE-CoMIMO interfaces.
- Control authority: A policy envelope lets the network set parameter ranges, resource budgets, and security policies while the hub makes fast local decisions within those limits.This division avoids slow per-decision network approval while constraining duplication, interference, and unfair resource use.
- Interoperability: Standardizing token headers and exchange procedures while leaving payload representations implementation-specific can support multi-vendor interoperability without constraining future encoders.
- Trust and privacy: Joint sensing and communication expose first-person video, sensing tokens, and device-state reports, while untrusted Co-UEs may inject false tokens or manipulate confidence.The paper therefore requires relay authentication, secure token exchange, user consent, malicious-device detection, and trust-aware recruitment.
- Validation: Validation must extend beyond simulation because performance depends on local geometry, body blockage, mobility, and hardware impairments.The proposed agenda includes OTA-calibrated models, measurement-based evaluation, multi-device experiments, conservative fallbacks, and limits on topology-change rates.
VI. CONCLUSION
Agentic UE-CoMIMO extends virtual antenna cooperation with user-side control over devices, roles, computation, and token budgets. Across two scenario studies, prediction, intent interpretation, and replanning improve sustainability and preserve continuity during assisting-device failures.
- Conclusion: Agentic UE-CoMIMO selects cooperating devices, assigns roles, chooses relay and traffic modes, places computation, and allocates token budgets according to intent and device state.
- Conclusion: The scenario studies distinguish cooperation, adaptive control, and agentic control as separate contributors to system performance.
- Conclusion: Prediction, intent interpretation, and feedback-driven replanning improve resource sustainability over longer sessions and maintain service continuity when an assisting device fails.