Source-linked AI summary

Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evidence

Ziheng Li, Xichen He, Haoyan Chen, Charlie Zou, Sheng Bai, Benjamin Yang, Mengyuan Wu, Jake Ledner, Yi-Jie Cheng, Akito Yamauchi, Dishita G Turakhia, Steven Feiner, Paul Sajda

arXiv:2608.30369v1cs.AIcs.HC

TL;DR

Attention-limited vigilance tasks require assistance when candidates exceed human inspection capacity and target classes can shift without warning. OLIVE adapts a frozen vision-language model online by fusing explicit actions with fixation-locked EEG while estimating source reliability. Across three studies, the full agent achieved the strongest convergence and within-session assistance outcomes, including faster recovery after silent target switches, though modest samples and an untested fixed prevalence anchor limit the claims.

  • Problem

    Operational vigilance tasks generate more candidates than users can inspect in time, while labels are sparse, asymmetric, heterogeneous, and must be processed online within task time budgets.

  • Method

    OLIVE maintains item-level target posteriors and jointly adapts a frozen VLM using behavioral actions and fixation-locked EEG while estimating each evidence source’s reliability online.

  • Results

    OLIVE-IE achieved the highest guidance convergence rate, 96.8% at 28.3 s, and no variant matched it on both convergence rate and speed.

  • Takeaways & Limitations

    The combined agent produced reliable within-session augmentation and faster readaptation when the target class changed silently, supporting passive neural copilots for attention-limited tasks.

  • Takeaways & Limitations

    US2 and US3 used modest samples of n=8–14 per condition, and OLIVE’s fixed prevalence anchor π remains untested under drifting prevalence.

Abstract

from arXiv · show

We present OLIVE, a framework for adapting a foundation model to provide real-time assistance in temporally demanding, high-stakes, and dynamic tasks. We show that passive EEG, fused online with behavioral evidence, can meaningfully extend the number of targets users detect and engage beyond their unaided action bandwidth. OLIVE learns from both explicit behavioral signals (the targets the user shoots down in an XR first-person shooter game) and implicit physiological signals (fixation-locked EEG) to provide timely guidance, continuously adapting a frozen vision-language model's inference on which items are task-relevant by jointly estimating per-source reliability without manual labels or offline training. Through three user studies, including two live deployments of an assistive agent driven by OLIVE in XR, we show that OLIVE Pareto-dominates prior test-time adaptation frameworks, achieving the highest convergence rate at comparable convergence speed. Combining implicit physiological and explicit behavioral signals, the OLIVE agent produces the largest and most reliable within-session improvement to a user's ability to detect and engage targets, largely independent of the individual's skill. When the target switches silently, the agent that uses both behavioral and physiological signals reconverges significantly faster than the behavior-only agent (1.27 times faster on average, p = .008), restoring trustworthy guidance at the moment the task changes, precisely when reliable assistance matters most.

1 Introduction

Operational vigilance tasks overload human attention because candidates exceed timely inspection capacity, target classes can shift silently, and actions expire. OLIVE addresses this gap by adapting a foundation model online from sparse behavior and noisy fixation-locked EEG.

  • Operational vigilance tasks involve more candidates than people can inspect before action windows expire, with target classes that may shift without warning.
  • Passive BCI signals can provide task-relevant physiological evidence during natural interaction without requiring explicit user commands.
  • Existing methods separately address unlabeled visual adaptation, reliability-weighted annotations, or offline brain-signal relevance, leaving online human-contingent assistance unresolved.
  • OLIVE fuses behavioral actions and fixation-locked EEG, estimates source reliability online, and adapts a frozen VLM within the task time budget.
  • Across three user studies, OLIVE evaluates convergence, live performance improvement, and rapid readaptation after silent target-class shifts.

2 Background

The background positions OLIVE at the intersection of online foundation-model adaptation, crowd-style reliability estimation, passive physiological evidence, and shared autonomy. Its distinguishing role is online EM adaptation from heterogeneous human-contingent evidence within seconds-long task windows.

  • Operational vigilance tasks combine continuous candidate streams, limited inspection capacity, expiring action windows, and silently shifting target classes.
  • Foundation-model test-time adaptation typically relies on unlabeled visual streams, while crowd-learning models do not accommodate heterogeneous evidence or VLM updates.
  • Reinforcement learning is poorly matched to seconds-long windows, and passive human–AI interaction differs from frameworks assuming deliberate discrete input.
  • OLIVE fills this gap by adapting a foundation model online from heterogeneous human-contingent evidence through an EM update within the task time budget.
  • Fixation-locked EEG can distinguish target from distractor fixations, but prior systems generally use decode-then-act pipelines rather than feeding physiological posteriors into online EM.
  • Effective assistance must match machine initiative to user capacity and confidence, while covert neural sensing raises requirements for accurate user mental models and autonomy.

3 Methods

OLIVE treats item targetness as a latent state and updates it online by combining visual, behavioral, and fixation-locked EEG evidence with learned source reliabilities. The method is instantiated in an XR shooter where guidance and user actions provide closed-loop evidence during each task round.

  • 3.1 Problem Statement: The problem has incomplete, asymmetric, heterogeneous labels: actions are sparse and consequential, while fixation-locked EEG arrives earlier but noisier.
  • 3.2 Task Environment: SpaceShooter models operational vigilance with mixed enemy and friendly ships, an initial target prevalence of π0 ≈0.30, and heavier penalties for shooting friendly ships.
  • 3.2 Task Environment: Round difficulty controls fleet size, while item beliefs reset and reliability parameters plus the VLM prompt warm-start from the previous round.
  • 3.3 Calibration: A visual-search calibration phase trains a participant-specific fixation-locked EEG decoder from labeled static-array fixations before SpaceShooter sessions.
  • 3.4 OLIVE: Online Latent Inference from Variable Evidence: OLIVE maintains item-level target posteriors and jointly updates evidence-source reliability and the VLM while treating channels as noisy rather than perfectly accurate.
  • 3.4 OLIVE: Online Latent Inference from Variable Evidence: Positive actions provide explicit target evidence, whereas non-actions remain unlabeled because users may not reach recognized targets under load.
  • 3.4 OLIVE: Online Latent Inference from Variable Evidence: The first fixation on each item contributes a soft EEG target probability, with online confidence and discriminativeness estimates down-weighting noisy physiological evidence.
  • 3.4 OLIVE: Online Latent Inference from Variable Evidence: The online EM loop updates whenever an action or fixation arrives, combining visual, action, and physiological log-odds before prevalence anchoring.

4 User Study 1 (US1): Convergence and Robustness

US1 evaluated OLIVE across evidence configurations and baselines using convergence rate and speed, with OLIVE-IE reaching the strongest combined operating point. Its guidance remained robust across user evidence quality and overload-related conditions.

  • RQ1.1 Pareto frontier: OLIVE-IE achieved the highest guidance convergence rate, 96.8% (92/95 rounds; 28.3 s), while no variant matched it on both convergence metrics.Its belief convergence was 78.9% at 37.3 s, placing it on the high-rate frontier.
  • RQ1.1 Pareto frontier: OLIVE-IE locked onto the correct target set within the first third of 90-second rounds in 97% of rounds.This provided trustworthy agent support for the remainder of those rounds.
  • RQ1.2 Sensitivity to evidence quality: OLIVE showed no sensitivity to shot accuracy, whereas olive-base-E’s belief convergence time increased with shot accuracy.OLIVE-IE guidance convergence rate also improved with higher EEG quality, supporting the contribution of the implicit channel.
  • RQ1.2 Sensitivity to evidence quality: With low EEG quality, OLIVE-IE maintained 94%–100% guidance convergence while OLIVE-I declined from 53% to 36%.The reliability update suppresses the physiological channel when its evidence becomes unreliable and falls back toward behavioral evidence.
  • RQ1.3 Overload robustness: Under cognitive overload, OLIVE-I took longer to converge as overwhelm increased, while olive-base-E showed lower belief convergence with higher workload.These modality-specific degradations motivated combining behavioral and physiological evidence.

5 User Study 2 (US2): Live Agent Performance

US2 evaluates OLIVE in live XR deployment, measuring whether convergence persists and whether guidance improves target throughput across repeated rounds. OLIVE-IE maintained near-total convergence, produced the only significant within-session gain, and improved similarly across skill levels.

  • 5.2 Agent Support Policies: Visual Guidance Cues.: OLIVE withheld guidance while beliefs were unstable, then activated directional lines, outline highlights, and aim assist according to belief strength.Hard cues targeted OLIVE’s top-4 items, while soft cues were graded continuously; Oracle marked all attacking targets.
  • 5.4.1 RQ2.1: OLIVE Guidance Convergence Is Maintained in Live Deployment.: 99.4% guidance convergence for OLIVE-IE in live deployment matched or exceeded its 96.8% offline benchmark despite longer convergence times.Belief convergence also improved from 78.9% offline to 84.6% live, while guidance convergence took 74.5 versus 28.3 seconds.
  • 5.4.2 RQ2.2: OLIVE-IE Produces Reliable Within-Session Throughput Improvement.: +0.031 kills/s was OLIVE-IE’s only significant within-session throughput improvement, with one-sample p = .003.OLIVE-E and Control showed marginal trends, while Oracle’s +0.024 kills/s improvement was not significant.
  • 5.4.2 RQ2.2: OLIVE-IE Produces Reliable Within-Session Throughput Improvement.: OLIVE-IE’s throughput gain was essentially independent of baseline skill, whereas OLIVE-E’s improvement concentrated among lower-skill operators.The reported skill pattern is consistent with OLIVE-IE compensating for what operators cannot cover unaided.
  • 5.4.3 Operator Reliance and Trust.: Reliance and trust increased from Control to OLIVE-E to OLIVE-IE to Oracle, with OLIVE-IE exceeding OLIVE-E on all three reliance measures in US3.In US2, Oracle significantly separated from Control on agent trust and gaze reliance.

6 User Study 3 (US3): Adapting to Silent Target Changes

US3 tests whether OLIVE can recover when the target definition changes silently during a round. OLIVE-IE achieved faster guidance reconvergence than behavior-only OLIVE-E and produced the largest within-session gain in new-target throughput.

  • 6.3 Results and Discussion: 100% switch detection across Control, OLIVE-E, OLIVE-IE, and Oracle shows performance differences reflected reorientation ability rather than unawareness.Switch detection rates were Control 100%, E 98.6%, IE 100%, and Oracle 100%.
  • 6.3.1 RQ3.1: OLIVE-IE Reconverges Faster After the Switch.: 14.7 s faster guidance reconvergence was achieved by OLIVE-IE than OLIVE-E after silent target switches (53.8 ± 4.1 s vs. 68.5 ± 3.8 s, p = .008).Both conditions achieved 100% guidance reconvergence; belief reconvergence favored OLIVE-IE in rate and mean time, but the time difference was not significant.
  • 6.3.2 RQ3.2: OLIVE-IE Produces Reliable Within-Session Readaptation.: OLIVE-IE showed the largest within-session new-target throughput improvement: Δ = +0.070 kills/s, p = .001.Oracle also improved significantly, whereas OLIVE-E and Control showed positive but marginal trends.
  • 6.3.2 RQ3.2: OLIVE-IE Produces Reliable Within-Session Readaptation.: OLIVE-IE’s readaptation gain was essentially skill-independent, while unaided Control’s gain was strongly skill-dependent.The reported correlations were r = −0.24, p = .57 for OLIVE-IE and r = +0.68, p = .01 for Control.
  • 6.3.3 RQ3.3: EEG Integration Produces a Larger Within-Session Gain.: OLIVE-IE’s within-session gain exceeded OLIVE-E’s, while OLIVE-E did not clearly separate from Control, supporting an EEG augmentation advantage.The paper attributes the advantage to a longer window of reliable guidance after each switch.

7 Discussion and Future Work

OLIVE’s discussion frames adaptive assistance as trust-calibrated, skill-aware, and dependent on deployment assumptions. The authors identify silent target switches, evidence fusion, and prevalence stability as important directions and boundaries.

  • Trust and human–AI teaming: Reliance on OLIVE-IE increased during sessions, especially after silent target switches, while reliance on behavior-only assistance declined.Gaze-reliance on the EEG-informed agent rose by +0.07 rating units per round (p < .05) in US3, whereas behavior-only reliance eroded.
  • Evidence fusion and skill: OLIVE-IE produced skill-independent gains, whereas OLIVE-E’s improvements concentrated among lower-skill operators.The evidence-fusion configuration therefore exposes a task-tunable precision–coverage trade-off.
  • Silent target switches: EEG helped the agent detect silent target switches independently of shot accuracy, while behavior-only adaptation depended on changing shot patterns.The compensatory OLIVE-E effect weakened in US3 (r = −0.11, ns) compared with US2 (r = −0.63).
  • Deployment boundaries: Deployment beyond the XR setting requires revalidating known target prevalence, fixation-to-item alignment, and a confirmation signal.These assumptions differ across radiology, air traffic control, and surveillance.
  • Future work: Future extensions include multi-fixation fusion and additional soft annotators such as dwell time, voice annotation, pupil dilation, or response-locked ERN.The current architecture used only the first fixation, although US1 reached 96.8% guidance convergence.
  • Limitations: The fixed prevalence anchor π remains untested under drifting prevalence, although it rescales posterior magnitudes without changing belief rankings.The authors identify online prevalence estimation as future work.

8 Conclusion

The conclusion presents OLIVE as a passive XR copilot that extends attention-limited performance while preserving operator agency. Its assistance combines minimal visual cues with broader perceptual coverage and evidence from user actions.

  • Conclusion: Targets and non-targets share motion dynamics, so classification must rely on semantic features and accumulated evidence rather than motion cues.This design prevents trajectory-based class identification.
  • Conclusion: Silent target switches make prior knowledge misleading by changing which class attacks without explicit announcement.Previously valid targets disengage while the new target class begins attacking.
  • Conclusion: The task uses forgiving homing assistance so misses more directly reflect decision errors than aiming difficulty.This preserves the interpretation of shots as meaningful behavioral evidence.
  • Conclusion: The environment constrains performance through competing ships, limited action windows, and a one-second weapon cooldown.The cooldown discourages mindless spraying and forces prioritization among perceived targets.
  • Conclusion: OLIVE’s minimal red outlines guide attention without automating the user’s decision.The representation is designed to preserve agency at the point of action.
  • Conclusion: The wingman agent extends the user’s effective perceptual field by scanning outside the current field of view.Edge-of-screen indicators direct the user toward highlighted targets outside their view.

A.9 Experimental Setup

The experimental system combines XR interaction, eye tracking, EEG processing, vision services, and online OLIVE adaptation through a low-latency runtime architecture. Its model fuses visual, behavioral, and physiological evidence while maintaining a prevalence-anchored posterior.

  • Hardware and monitoring: The setup used a Varjo XR-3 headset with 200 Hz eye tracking and a 20-electrode, 256 Hz wireless EEG system.The live interface exposed XR, eye-tracking, and EEG streams for monitoring without interrupting performance.
  • Runtime architecture: Four communicating processes connect Unity with Vision, OLIVE, and Physio services through gRPC.The Vision Service routes crops and shot events to OLIVE while forwarding fixation events to Physio.
  • Runtime architecture: Unity captures user and agent camera streams, sends shot events, and renders OLIVE’s current beliefs as guidance cues.Shot identities become explicit positive labels for the OLIVE Service.
  • Evidence model: Fixation-locked EEG supplies one soft physiological label per item, with decoder quality represented by a per-fixation weight.Only the first fixation contributes physiological evidence.
  • Evidence model: The visual channel adapts only a virtual prompt while keeping the CLIP backbone frozen.Quality-weighted item crops from user and agent cameras contribute to the visual evidence.
  • Evidence model: Behavioral evidence labels fired-on ships as positive examples, while non-actions remain unlabeled.This avoids treating missed opportunities as evidence that an item is irrelevant.
  • Online EM: A global bias shifts item log-odds so the mean posterior matches prevalence π without changing belief rankings.This prevents zero collapse when distractor evidence dominates and preserves the ordering of cued items.

C.3 Parameterization and robustness

OLIVE parameterizes online evidence reliability and guidance from behavioral and fixation-locked EEG signals, while the calibration pipeline trains participant-specific physiological decoders. Robust posterior P300 discrimination supports the EEG channel, whereas pupil signals were excluded because rapid luminance changes confounded them.

  • Reliability parameterization: OLIVE estimates action and physiological reliability parameters online, down-weighting evidence sources when their observed signals become unreliable.The action parameters model false positives and true positives, while physiological class means characterize EEG evidence.
  • Physiological evidence: The system requires a calibrated per-fixation physiological probability, but remains decoder-agnostic about how that probability is produced.Any module returning a calibrated target-presence probability can serve as the implicit evidence source.
  • Calibration pipeline: Participant-specific calibration uses fixation-locked EEG epochs labeled by the ground-truth identity of each fixated item.Fixations shorter than 100 ms are excluded, and each epoch spans −100 to 800 ms relative to fixation onset.
  • Calibration pipeline: The decoder passes the full preprocessed 20-channel fixation epoch to a learned spatiotemporal model rather than using hand-engineered features.The design is motivated by N2pc and P300 components associated with attentional selection and target confirmation.
  • Robustness and scope: Pupil data were excluded because rapid luminance transients confounded cognitive responses and the slower TEPR window was difficult to align with individual fixations.The paper treats this as a practical scope decision rather than evidence that TEPR is intrinsically uninformative.
  • Physiological evidence: A posterior P300 at 300–600 ms and posterior sites provides robust target–non-target discrimination during calibration.The strongest discrimination appears at Pz, P3, P4, POz, O1, and O2, while the target–distractor difference is minimal in early windows.

E.8 Error-Related Negativity During Friendly-Fire Events

The friendly-fire analysis examines response-locked EEG errors and related decoder and portable-form-factor analyses. Error-related activity was visible but nonsignificant, while simulated portable EEG performance favored the OLIVE strip over Emotiv EPOC X with important simulation caveats.

  • Error-related negativity: Friendly-fire shots provide objectively incorrect, score-penalized response errors for comparing response-locked EEG with correct enemy-destruction trials.The analysis uses response-locked epochs from 38 US1 participants with sufficient clean error epochs.
  • Error-related negativity: 1.05 μV error-minus-correct positivity in the 40–140 ms window did not reach conventional significance (p = .096).The reported statistic was t(37) = 1.71, and cluster permutation found no significant cluster.
  • Error-related negativity: The frontocentral ERN/Pe pattern was attenuated under heterogeneous, motion-prone, high-workload VR conditions.The paper attributes the limited effect to intentional versus accidental errors, headset artifacts, and competition with ongoing perceptual demands.
  • Portable EEG: The decoder’s channel-ablation analysis identifies POz, Cz, Fp2, and T3 as the four most important channels, although individual drops remain modest.The joint multi-channel tokenizer means no single electrode has an independent representational role.
  • Portable EEG: The OLIVE strip significantly outperformed Emotiv EPOC X by 0.02 AUC in simulated paired testing (t(40) = 1.72, p = .047).Group means ranged from 0.664 for EPOC X to 0.685 for the OLIVE strip; no other pairwise comparison was significant.
  • Portable EEG: Portable-form-factor AUCs are simulated priors rather than measurements from actual devices, and interpolation can introduce approximation error.The paper cautions that native device recordings or retrained decoders could differ substantially, especially for Muse 2 extrapolation.

F Convergence Metrics

The paper defines convergence as stable, trustworthy guidance rather than internal model accuracy alone. Its metrics combine top-4 ranking stability with correctness, and its evidence configurations separate behavioral, physiological, and joint assistance.

  • Convergence rationale: Convergence matters because an unconverged agent can actively mislead users with plausible-looking distractors.The paper therefore defines convergence by when guidance becomes reliably worth acting on.
  • Convergence metrics: Top-4 ranking stability uses Jaccard similarity between the current and previous sets of four highest-posterior items.High similarity indicates consistent surfaced items; low similarity indicates rapidly changing suggestions.
  • Convergence metrics: Belief convergence requires AUC > 0.90, top-4 stability S_t ≥ 0.75, and both conditions sustained for at least 10 seconds.This criterion captures a stable and correct posterior-based ranking.
  • Convergence metrics: Guidance convergence requires Precision@4 = 1.0 and S_t ≥ 0.75 sustained for at least 10 seconds.All four highlighted items must be true targets, preventing stable-but-wrong guidance from qualifying.
  • Evidence configurations: The E, I, and IE configurations isolate explicit behavior, implicit EEG, and their combined evidence, with IE as the primary condition.The baselines additionally test frozen reliability weights, prompt adaptation, and feature-cache retrieval under matched evidence streams.
  • Task characterization: Adaptive difficulty preserves hit rate near 0.67–0.71 while mothership health falls from 85.7 at difficulty 1 to 52.6 at difficulty 5.The results indicate increasing scene pressure while participants maintain comparable shot accuracy.
  • Task characterization: Participants varied substantially in adaptive difficulty: some reached level 5, others plateaued at level 2, and the median maximum was 3 (IQR: 3–4).Most participants stabilized within 4–6 rounds, with later fluctuations of approximately ±1 level.
  • Subjective ratings: Mothership health correlates negatively with overwhelm (r = −0.31) and workload (r = −0.22), while hit rate correlates positively with confidence (r = 0.43).Confidence also increases with difficulty, which the paper interprets as a selection effect rather than a causal difficulty effect.

G.5 Simulation-based error bound for the 16-round QUEST+-style calibration

The paper evaluates whether a 16-round QUEST+-style calibration can recover an individualized challenging-but-playable difficulty. In simulation, estimation error was usually within half a difficulty level, supporting its use for initializing later sessions.

  • Motivation: The calibration concern is whether only 16 adaptive rounds can recover a participant-specific operating point accurately enough for initialization.The full procedure contains two fixed practice rounds followed by 16 adaptive rounds.
  • Simulation procedure: The simulation models 2,500 virtual participants with ability uniformly sampled from 1 to 6 and slopes sampled from five values between 0.4 and 1.2.Scores include logistic noise, clamping to 0–100, and five discretized outcome categories.
  • Simulation procedure: The adaptive policy selects each next difficulty to maximize expected information gain, defining d★ where expected score equals 70.The target is the paper’s “challenging-but-playable” operating point.
  • Results: Median absolute estimation error was 0.20 difficulty levels, with 90.0% of simulated participants within ±0.5 and 99.5% within ±1.0.Mean signed error was 0.006, indicating negligible bias in the simulation.
  • Results: The simulated calibration was considered sufficient to initialize later sessions within an approximately half-level individualized difficulty band.This conclusion is based on the modeled score-generation process rather than direct measurements from participants.

H.3 Adaptive Difficulty Trajectories

US2 adaptive difficulty was balanced across conditions, while OLIVE achieved high live convergence. In US3, comparable task demands support interpreting throughput differences as adaptation performance rather than difficulty confounds.

  • US2 adaptive difficulty: All conditions started near difficulty 3–4 and converged to difficulty 5 by mid-session in US2.Stratified assignment balanced starting ability across conditions.
  • US2 adaptive difficulty: OLIVE-E and Oracle more reliably reached the difficulty ceiling, with every participant reaching difficulty 6 versus Control’s mean of 5.38 ± 0.86.The overall condition difference was not significant: F=1.50, p=.239.
  • OLIVE convergence: Live convergence reached 99.4% for OLIVE-IE and 97.1% for OLIVE-E, exceeding US1 offline benchmarks of 96.8% and 79.0%, respectively.The comparison concerns near-total guidance convergence across the IE and E conditions.
  • US3 adaptive difficulty: In US3, all conditions reached difficulty 4–5 by mid-session and sustained similar trajectories through data-collection rounds.The session included a randomized silent target switch unknown to participants.
  • US3 adaptive difficulty: US3 showed no significant between-condition difficulty differences: F=0.20, p=.896, with all pairwise p>.40 and |Δ|≤0.38.The authors conclude that conditions faced comparably challenging enemy compositions when silent switches occurred.

I.4 Subjective Measures

Subjective measures indicate that EEG augmentation was associated with lower temporal demand than behavior-only assistance, while automatic Oracle adaptation was associated with elevated frustration. Participants also rated EEG-assisted adaptation speed higher.

  • Wingman adaptation: OLIVE-IE participants rated wingman adaptation speed higher than OLIVE-E participants, 5.1 versus 4.4 on a 1–7 scale.The subjective difference was consistent with faster objective reconvergence for OLIVE-IE.
  • NASA-TLX: NASA-TLX profiles were collected after the experiment using Welch t-tests across conditions, with sample sizes of 7, 8, 8, and 5 for Control, E, IE, and Oracle.The profile used an adapted 1–7 scale rather than the conventional 0–100 format.
  • NASA-TLX: Temporal demand was higher for OLIVE-E (M=5.2) and Oracle (M=5.33) than Control (M=2.8), while OLIVE-IE was intermediate at M=4.0.The measures used an adapted 1–7 response scale.
  • NASA-TLX: Frustration was low and comparable across Control, OLIVE-E, and OLIVE-IE (M=3.2–3.5), but elevated for Oracle at M=5.33.The Oracle-versus-Control comparison was not statistically significant: t=2.02, p=.118.
Loading 2608.30369v1…