Source-linked AI summary
Cross-Modal Guidance for Out-of-View Object Search in Simulated Prosthetic Vision
Adyah Rastogi, Apurv Varshney, Tobias Höllerer, Michael Beyeler
TL;DR
Out-of-view guidance may work differently when simulated prosthetic vision severely limits visual bandwidth and guidance must share the sparse scene representation. The paper compares visual, haptic, and auditory cues carrying the same horizontal target-offset information across two SPV conditions and finds that tested nonvisual cues were faster overall and during acquisition, with effects varying by search stage and visual constraint.
Problem
Whether visual, auditory, and haptic guidance retain the same tradeoffs under prosthetic vision, where guidance and scene content share a sparse visual representation, is unknown.
Method
Nineteen sighted participants performed object search under 10 × 10 and 20 × 20 simulated prosthetic vision using visual, haptic, or auditory cues encoding the same horizontal target-offset variable.
Results
All three modalities reduced search time and head movement, while tested auditory and haptic cues produced approximately 25% faster overall search than visual guidance and also shortened post-acquisition search.
Takeaways & Limitations
Under severe visual constraints, guidance performance depended on cue implementation and search stage, with tighter final alignment and improved uncued elevation localization in the 10 × 10 condition.
Takeaways & Limitations
The study used sighted participants experiencing SPV in a short session, so it does not reproduce blindness, long-term adaptation, or individual implant percepts; clinical generalization requires visual-prosthesis users.
Abstract
from arXiv · showhide
Out-of-view guidance is well established in virtual and augmented reality, but its effectiveness may depend on the visual bandwidth available to the user. We test this under simulated prosthetic vision (SPV), where visual guidance must share the same sparse representation used to inspect the scene. Nineteen participants performed object search under two SPV conditions differing in electrode density and phosphene spread (10x10 and 20x20) and four guidance conditions (no guidance, visual, haptic, audio) all driven by the same horizontal target-offset variable. All three modalities reduced search time and head movement. The tested auditory and haptic cues produced approximately 25% faster overall search and 11-13% faster target acquisition than the visual cue, despite similarly direct orienting trajectories. The tested haptic and auditory cues also shortened post-acquisition search. Final head-target angular offset was reduced substantially more in the 10x10 SPV condition; there, all three cues also reduced vertical localization error by approximately 45-58% despite providing no elevation information. Under severe visual constraints, guidance performance depended on cue implementation and search stage.
1 INTRODUCTION
The paper examines whether visual, auditory, and haptic cues can guide object search when simulated prosthetic vision severely limits the visual channel. In a controlled comparison using the same horizontal target-offset information, nonvisual cues outperformed the tested visual cue and benefits varied across search stages and visual conditions.
- Motivation: Severely bandwidth-limited prosthetic vision requires users to scan environments because relevant camera content may lie outside the represented field of view.This mismatch makes out-of-view guidance important for object search, an accessibility problem for people who are blind.
- Approach: The study compared visual, auditory, and haptic cues encoding the same horizontal target-offset variable under simulated prosthetic vision.The cues provided neither target elevation nor identity, and modality-specific mappings were used without equating physical cue parameters.
- Findings: All three modalities reduced search time and head movement, while auditory and haptic guidance produced approximately 25% faster overall search than visual guidance.The study used 19 sighted participants across 10 × 10 and 20 × 20 biologically motivated SPV conditions.
- Findings: Guidance benefited both target acquisition and subsequent completion, with the tested haptic and auditory cues shortening post-acquisition search.The analysis explicitly partitioned search into target acquisition and the interval to response.
- Findings: The tested auditory and haptic cues outperformed visual guidance under severe visual constraints, while the 10 × 10 condition produced tighter final alignment and improved uncued elevation localization.These results indicate that cue implementation and search stage shaped guidance performance.
2 RELATED WORK
Prior work improves either the information represented within prosthetic vision or the user’s access to out-of-view content, but cross-modal guidance remains underexplored when cues share a sparse visual representation. This paper frames prosthetic vision as a stringent test of modality and cue-design tradeoffs.
- Object Search and Prosthetic Vision: Assistive object search requires both spatially locating a target and obtaining enough perceptual information to identify it under a limited visual channel.Prosthetic vision compounds these demands through small field of view and low spatial resolution.
- Object Search and Prosthetic Vision: Prior prosthetic-vision approaches improve represented scene information through image simplification, semantic segmentation, and depth- or task-dependent filtering.These methods address content within the prosthetic field of view rather than initially bringing relevant content into it.
- Prior Guidance Studies: Earlier studies reported benefits from visual saliency cues for SPV search and from auditory and haptic feedback for navigation, without directly comparing these modalities for the same search guidance task.The present comparison addresses this gap.
- SPV Modeling: The axon-map model represents phosphene shape using axonal and radial spread parameters rather than regular independent points of light.BionicVisionXR uses this psychophysically validated model for head-directed SPV tasks.
- Out-of-View Guidance: Traditional out-of-view techniques encode direction or distance with display-boundary, arrow, radar, or focus-plus-context indicators in VR and AR.Comparative work shows that guidance design affects acquisition time, head rotation, and situation awareness under restricted fields of view.
- Cross-Modal Guidance Under Visual Constraints: Under prosthetic vision, visual guidance must survive the same severe spatial bottleneck as scene content, whereas auditory and haptic cues can redirect attention without adding visual content.The paper therefore compares modality-specific implementations encoding the same horizontal target-offset variable before and after acquisition.
3 METHODS
The experiment recruited 19 sighted adults with normal or corrected-to-normal vision and included all participants in the analysis under institutional ethical approval.
- Participants: Nineteen participants aged 19–27 years were recruited from the university community, including 11 female and 8 male participants.All reported normal or corrected-to-normal vision and no known visual or neurological impairments.
- Participants: All participants completed the experiment and were included in the analysis after providing written informed consent.The study was approved by the university IRB, and participants received course credit or monetary compensation.
3.2 Apparatus
The apparatus used an immersive, head-tracked VR setup that allowed seated participants to rotate their heads and upper bodies while recording head pose and controller state.
- Hardware and Tracking: The experiment ran in Unity on an HTC VIVE Pro Eye headset with six-degree-of-freedom head tracking.Participants sat in a swivel chair and could rotate their head and upper body freely.
- Hardware and Tracking: Tracked VIVE controllers collected responses and delivered vibrotactile feedback in the haptic condition.Head pose and controller state were recorded continuously.
3.3 Simulated Prosthetic Vision
The study simulated prosthetic vision with an axon-map model that converts electrode activation into a sparse phosphene percept. It compared 10×10 and 20×20 electrode arrays spanning the same retinal area and visual field but differing in electrode density and phosphene spread.
- Multiple activated electrodes produce a spatial pattern of light spots or streaks called the prosthetic percept.
- The axon-map model predicts each frame’s percept from electrode activations, with ρ controlling radial phosphene spread and λ controlling nerve-fiber-aligned elongation.The VR camera image was sampled at simulated electrode locations to determine activation.
- 10×10 and 20×20 arrays covered the same 4.5×4.5 mm retinal area and 30°×30° visual field but used different electrode spacings and radial spread values.The 10×10 condition used 500 µm spacing and ρ = 400 µm; the 20×20 condition used approximately 237 µm spacing and ρ = 200 µm.
- λ was fixed at 100 µm in both SPV conditions to limit axon-aligned elongation and hold phosphene shape constant.
- The two conditions represented substantially different prosthetic-resolution regimes rather than small display-resolution changes.The electrode count differed fourfold, a scale described as substantial in visual prosthesis design.
- The percept-rendering pipeline normalized image intensity, sampled electrode locations, scaled activation, and executed the axon-map model on the GPU in Unity.The exported pipeline was verified against the native PyTorch implementation using identical inputs and experimental parameters.
3.4 Task and Environment
Participants searched for specified objects in a virtual workspace viewed exclusively through SPV. Targets appeared at wide-ranging desk locations, often outside the horizontal prosthetic field of view, requiring head rotation and subsequent SPV-based localization.
- The virtual environment contained a semicircular desk surrounding the seated participant.
- The target set contained 14 familiar desk and household objects varying in size and coarse shape.
- Objects appeared at eight predefined desk locations spanning approximately 200° of horizontal azimuth, from −100° to +100°.
- Trials included single-target scenes and cluttered scenes containing the target with four distractors.
- Camera height varied by −0.5, 0, +0.5, or +1.0 m across trials, requiring vertical localization because guidance supplied only horizontal information.
- Targets typically began outside the 30°×30° horizontal prosthetic field of view, so participants rotated their heads to bring them into the represented region.
- Azimuthal acquisition was the first frame in which the target center entered the ±15° horizontal prosthetic field of view, independent of object size.The target could later leave and re-enter the field while participants localized it vertically and distinguished it from distractors.
3.5 Guidance Conditions
All guidance conditions encoded the same horizontal target-offset variable through modality-specific cues. Visual cues shared the SPV display, whereas audio and haptic cues conveyed direction and urgency outside the visual representation.
- Cue laterality indicated whether the target was left or right, while cue urgency increased as the absolute horizontal offset decreased.
- Within ±5° of horizontal alignment, all modalities switched to a centered state without providing target elevation information.Participants therefore still relied on SPV to complete the search.
- The baseline condition provided no directional guidance.
- Visual guidance used pulsating cues at the left and right edges of the prosthetic display, rendered through the same SPV pipeline as the scene.One or both edge cues were active depending on whether the target was outside or within ±5° of alignment.
- Audio guidance used stereo-panned beeps for target direction and varied the inter-beep interval from 0.05 s near alignment to 1.0 s at offsets of 135° or greater.
- Haptic guidance used direction-matched vibrotactile pulses, with both controllers vibrating within ±5° of alignment.Inter-pulse interval, intensity, and duration varied across the 0–135° target-offset range.
3.6 Experimental Design and Procedure
The experiment used a within-subjects design crossing four guidance conditions with two SPV conditions and two clutter levels. Participants completed 160 trials across balanced and counterbalanced blocks, with practice and post-block assessments.
- The within-subjects design crossed guidance, SPV condition, and scene clutter in a 4×2×2 arrangement.Guidance included baseline, visual, haptic, and audio; SPV included 20×20 and 10×10; clutter included one versus five objects.
- Each participant completed 160 experimental trials in eight 20-trial guidance-by-SPV blocks.
- Every block contained 10 single-object and 10 cluttered trials drawn from predefined target and distractor configurations.
- Guidance order followed a balanced Latin square, while SPV order was independently counterbalanced across participants.Ten participants completed 20×20 first and nine completed 10×10 first.
- Participants practiced the cues under normal vision before searching for text-named targets through the assigned SPV and guidance condition.They searched by rotating their head and upper body and were instructed to center the target in the prosthetic field of view.
- After each block, participants rated search difficulty on a 1–10 scale and later identified the most and least helpful guidance methods.
3.7 Measures
The study measured search time, acquisition and post-acquisition intervals, head-target alignment, head rotation, excess yaw, response alignment, and subjective difficulty during object search.
- Search time measured the interval from scene onset to participant response.
- Azimuthal acquisition latency measured time until the target center first entered the ±15° horizontal prosthetic FOV, marking represented-region entry rather than response criterion.
- Post-acquisition elapsed time partitioned search time as Tpost = Tresponse − Tacq, from first horizontal-FOV entry to response.
- Final head-target angular offset measured three-dimensional separation between head-forward direction and the target direction at response.
- Excess yaw quantified accumulated absolute yaw beyond the minimum rotation required to bring the target center to the horizontal SPV boundary.
- Participants rated search difficulty on a 10-point scale and ranked modalities by helpfulness after completing blocks and the experiment.
3.8 Statistical Analysis
Trial-level outcomes were analyzed with mixed-effects models incorporating guidance, SPV condition, clutter, covariates, and participant-level random effects, with planned multiplicity-adjusted contrasts and sensitivity analyses.
- All 19 participants completed 160 trials each, yielding 3,040 trials with no outlier exclusions.
- Linear mixed-effects models included guidance, SPV condition, clutter, their interactions, standardized initial eccentricity, and trial order.
- Models used participant, participant-by-block, and where supported target-object random intercepts, fitted by restricted maximum likelihood.
- Planned contrasts compared guidance with baseline within SPV conditions, pooled nonvisual with visual guidance, and guidance effects across SPV conditions.
- Trajectory analyses used initially out-of-view trials with observed target-center entry, totaling n = 2,870, and replaced eccentricity with standardized initial out-of-view distance.
- Sensitivity analyses modeled censored acquisition times with a log-normal accelerated failure-time model and assessed acquisition-before-response using participant-clustered binomial GEE.
4 RESULTS
Guidance reduced search time, acquisition latency, head movement, and alignment error under both SPV resolutions. Nonvisual cues were faster than visual guidance, while coarse SPV showed especially large alignment and elevation benefits.
- 4.1 Effects of Guidance on Search Performance: All three guidance modalities reduced search time; pooled auditory and haptic guidance produced approximately 25% faster search than visual guidance in both SPV conditions.At 20 × 20, reductions versus baseline were 24% visual, 45% haptic, and 42% audio; at 10 × 10, they were 35%, 50%, and 52%.
- 4.2 Target Acquisition and Scanning Behavior: 12.7% faster acquisition occurred with nonvisual versus visual guidance at 20 × 20, and 11.3% faster at 10 × 10.All three modalities were faster than baseline in both SPV conditions.
- 4.2 Target Acquisition and Scanning Behavior: 98.1% of trials began with the target center outside the horizontal prosthetic FOV, and guided trials more often acquired the target before response than baseline trials.Acquisition before response occurred on 90.1% of baseline trials versus 97.6–99.1% of guided trials.
- 4.2 Target Acquisition and Scanning Behavior: Guidance increased full-FOV target alignment at response from 87.9% to 93.7–96.6% at 20 × 20 and from 65.3% to 86.1–91.6% at 10 × 10.
- 4.2 Target Acquisition and Scanning Behavior: 67–81% reductions in excess yaw and 37–56% reductions in total head rotation indicated more direct, less exploratory orienting across modalities and SPV conditions.
- 4.3 Post-Acquisition Search: Haptic and audio guidance shortened post-acquisition time, whereas visual guidance had little effect and did not retain a reliable benefit in the stricter subset.At 10 × 10, reductions were 51% haptic, 44% audio, and 23% visual.
- 4.4 Effects of Guidance on Final Head-Target Alignment: 52–61% reductions in final head-target offset occurred at 10 × 10, compared with 15–26% reductions at 20 × 20; every modality showed a larger effect under coarse SPV.
- 4.4 Effects of Guidance on Final Head-Target Alignment: At 10 × 10, guidance reduced elevation error by approximately 45–58% despite providing only horizontal target information.At 20 × 20, guidance did not reliably improve elevation error.
5 DISCUSSION
Under severe visual constraints, guidance improved both initial orienting and later search, but benefits depended on cue implementation and SPV condition. Auditory and haptic cues outperformed visual guidance in the tested implementations, while the study’s simulated, simplified task limits generalization.
- 5 DISCUSSION: All three guidance modalities accelerated target acquisition, reduced unnecessary head rotation, and improved performance after acquisition.The stage split separates initial orienting from subsequent target maintenance and search.
- 5 DISCUSSION: All three modalities removed most excess yaw and produced more direct convergence, indicating less exploratory scanning rather than merely shorter trials.This trajectory pattern explains the reduction in total head rotation.
- 5 DISCUSSION: Haptic and auditory guidance substantially reduced post-acquisition elapsed time in both SPV conditions, whereas the visual effect was less robust.At 20 × 20, only haptic and auditory guidance reliably reduced this stage; at 10 × 10, all three cues shortened it.
- 5.2 Cross-Modal Guidance Under Severe Visual Constraints: 11–13% faster acquisition and approximately 25% faster overall search occurred with auditory and haptic guidance than with visual guidance, despite similarly direct trajectories.The nonvisual advantage generalized across both 10 × 10 and 20 × 20 SPV conditions.
- 5.2 Cross-Modal Guidance Under Severe Visual Constraints: The observed ordering compares practical cue implementations rather than intrinsic sensory modalities because mappings were modality-specific and not physically matched.Differences in salience, discriminability, or sensorimotor mapping could contribute to the ordering.
- 5.3 Guidance Across Search Stages: The findings support cross-modal guidance when visual content is task-relevant and suggest adapting cues across acquisition and post-acquisition stages.The study also used sighted participants in short SPV sessions, so clinical generalization requires validation with visual-prosthesis users.
- 5.3 Guidance Across Search Stages: Guidance conveyed only horizontal target offset yet reduced final head-target offset and, in 10 × 10 SPV, improved uncued elevation localization.At 10 × 10, final head-target offset fell by 52–61%, showing that coarse directional information could support broader alignment.
- 5.4 Limitations and Future Work: The task assumed a stationary, perfectly known target in a seated workspace with only horizontal guidance, leaving more realistic spatial conditions for future testing.Uncertain detections, occlusion, moving or absent targets, and locomotion were outside the tested setting.