Source-linked AI summary

LandmarkLens: Predicting and Presenting Effective Landmarks for Mixed-Reality Urban Exploration

Chu Li, Yotam Sechayk, Jared Hwang, Jon E. Froehlich, Takeo Igarashi

arXiv:2608.30142v1cs.HC

TL;DR

People with poor SOD face challenges building cognitive maps, while conventional navigation tools prioritize efficient guidance over attention to environmental features. The paper studies gaze and landmark-description differences, then develops LandmarkLens, an MR system that uses a VLM to highlight relevant landmarks. In a preliminary evaluation with eight poor-SOD participants, LandmarkLens improved scene recognition and nearly doubled landmark recall in map sketches relative to a turn-by-turn baseline.

  • Problem

    Existing navigation tools prioritize efficient point-to-point guidance, and whether directing poor navigators toward effective landmarks facilitates cognitive mapping remained untested.

  • Method

    The paper combines a 20-participant VR landmark-attention study with LandmarkLens, an MR system that uses skilled navigators’ gaze and verbal strategies to guide VLM-based landmark highlighting.

  • Results

    LandmarkLens significantly improved scene recognition accuracy and nearly doubled landmark recall in map sketches compared with a turn-by-turn baseline among eight poor-SOD participants.

  • Takeaways & Limitations

    Guided attention can support landmark-level spatial knowledge acquisition for people with poor SOD, while broader route and survey learning remains less established.

  • Takeaways & Limitations

    The final evaluation involved eight participants and routes only in Tokyo, limiting established generalizability across populations and urban settings.

Abstract

from arXiv · show

People with a poor sense of direction (SOD) struggle to build cognitive maps for effective spatial navigation, and existing navigation tools prioritize efficiency over spatial learning. To understand how navigation strategies differ by ability, we conducted a landmark attention study with 20 participants (ten good SOD, ten poor SOD) who navigated across four Tokyo neighborhoods in virtual reality (VR). We found systematic group differences in both gaze behavior and the types of landmarks they verbally identify as effective. Based on these findings, we built LandmarkLens, a mixed-reality (MR) navigation system that uses a vision-language model (VLM) to identify and highlight navigation-relevant landmarks. A follow-up study with eight poor-SOD participants showed improved performance in scene recognition, suggesting that guided landmark attention can support landmark-level spatial knowledge acquisition for people with poor SOD, a first step toward broader spatial learning.

1 Introduction

The paper examines how navigation ability shapes landmark attention and introduces LandmarkLens to guide people with poor SOD toward effective landmarks during MR navigation.

  • Motivation: People with poor SOD experience stress navigating unfamiliar environments, while efficiency-focused turn-by-turn tools can divert attention from environmental features and spatial learning.These tools may also undermine independent navigation over time.
  • Motivation: Prior research links cognitive mapping to landmarks and reports that good navigators select fewer but more effective landmarks.The paper uses SBSOD thresholds of ≥5.0 for good navigators and ≤3.5 for poor navigators.
  • Research questions: The study asks how good and poor navigators’ gaze patterns differ and whether directing poor navigators toward effective landmarks improves spatial knowledge acquisition.Earlier work had characterized these differences primarily through verbal self-reports.
  • Landmark attention study: 20 participants navigated four Tokyo neighborhoods in VR while researchers collected concurrent eye-tracking and verbal descriptions of navigation-relevant features.This combined what participants said they noticed with what they actually looked at.
  • Landmark attention study: Good navigators scanned a broader horizontal scene extent and referenced durable cues, whereas poor navigators more often referenced transient and incidental features.Examples of durable cues included signage and permanent store markers; transient cues included campaign posters, parked bikes, and decorative details.
  • LandmarkLens: LandmarkLens uses a VLM-based pipeline and gaze-interactive MR overlays to highlight navigation-relevant landmarks alongside conventional turn-by-turn directions.The system translates skilled navigators’ gaze and verbal landmark strategies into landmark prediction and presentation.
  • Evaluation: Eight poor-SOD participants showed higher scene recognition accuracy with LandmarkLens, nearly twice as many recalled landmarks in map sketches, and comparable route-choice and map-structural accuracy.The evaluation compared LandmarkLens with a counterbalanced turn-by-turn baseline across two routes.
  • Contributions: The paper contributes a 20-person landmark attention study, LandmarkLens, a preliminary eight-person evaluation, and an open-source dataset of over 3,000 8K 360° street-view images.The dataset includes gaze and landmark annotations across Tokyo.

2 Related Work

Related work frames navigation as progression from landmark to route to survey knowledge and surveys landmark-enhancement systems. LandmarkLens differs by selecting landmarks from cognitive and perceptual differences between good and poor navigators.

  • Landmarks, routes, and cognitive maps: Spatial cognition research distinguishes landmark, route, and survey knowledge as progressively broader forms of environmental knowledge.Landmark knowledge concerns recognizing objects or scenes; route knowledge encodes landmark sequences and triggered actions; survey knowledge builds a cognitive map.
  • Landmarks, routes, and cognitive maps: Landmarks are stationary, distinct, salient reference points, including global landmarks for broad orientation and local landmarks for close-range decisions.Route and survey knowledge require integrating landmarks with spatial relationships beyond a single viewpoint.
  • Landmarks, routes, and cognitive maps: LandmarkLens evaluates all three knowledge levels through scene recognition, route choice, and map sketching.These tasks respectively measure landmark, route, and survey knowledge.
  • Individual differences in cognitive mapping: The SBSOD is a widely used 7-point self-report instrument for measuring individual differences in navigational ability.A crosscultural validation reported a population mean of 4.2 with SD = 1.1 for N = 550.
  • Individual differences in cognitive mapping: Prior interventions include repeated route exposure, practice with feedback, and verbalization, while good-SOD people select fewer effective landmarks and poor-SOD people select unreliable features.Whether instructing attention toward effective landmarks facilitates cognitive mapping remained untested before this work.
  • Landmark enhancements in MR: AR and landmark-based systems have supported wayfinding, mental-map development, semantic labeling, older adults, and users with low vision.Prior systems include virtual landmark maps, destination overlays, and researcher-curated salient-landmark detectors.
  • Landmark enhancements in MR: Existing systems select landmarks using general saliency, user familiarity, formative studies, or computational detection, but not cognitive and perceptual differences between navigator groups.LandmarkLens grounds selection in skilled navigators’ attention patterns to guide people with poor SOD.

3 Landmark Attention Study

The landmark attention study combined eye tracking and verbal descriptions during VR navigation to compare good- and poor-SOD participants. Good navigators surveyed broader scene regions, identified more durable wayfinding cues, and chose correct retrace directions more often.

  • Study Design: 20 participants, split evenly by SOD, navigated 360° street-view routes across four Tokyo neighborhoods while providing concurrent gaze and verbal data.The study used eye tracking and verbal descriptions to examine both attended and articulated landmarks.
  • Gaze Behavior: Good navigators exhibited greater total and horizontal gaze spread, whereas vertical spread did not differ significantly between groups.Total gaze spread was M_good = 0.13 versus M_poor = 0.11; horizontal spread differed significantly, while vertical spread did not.
  • Gaze Behavior: Good navigators showed greater head yaw and pitch variability, while fixation rate and duration did not differ significantly.Head yaw variability was M_good = 36.0° versus M_poor = 28.2°, and pitch variability was 10.7° versus 9.0°.
  • Navigation Performance: Good navigators chose the correct direction at retrace intersections more often than poor navigators: 75.5% versus 57.5%.The group effect on intersection accuracy was statistically significant, while retrace duration did not differ significantly.
  • Verbal Landmark Strategies: Good navigators mentioned wayfinding signage more often, whereas poor navigators more frequently referenced decorative details, building parts, pedestrians, bikes, and generic signage.Good navigators also attended to permanent store signs and the Tokyo Skytree, while poor navigators focused more on menus and temporary signage.
  • Verbal Landmark Strategies: When both groups viewed the same campaign poster, good-SOD participants rejected it as temporary while poor-SOD participants considered its possible persistence.This contrast shows that verbal landmark usefulness was evaluated differently even when gaze overlapped.

4 LandmarkLens

LandmarkLens converts gaze and verbal strategies from good navigators into an MR landmark-guidance pipeline. It crops panoramas, detects and filters landmarks with a VLM, removes duplicates, validates detections, and presents gaze-interactive overlays.

  • System Overview: LandmarkLens combines a landmark prediction pipeline with an MR interface that overlays navigation-relevant landmarks on conventional turn-by-turn directions.The overlays guide attention toward visual features selected as effective clues by skilled navigators.
  • Pipeline: The pipeline uses gaze-informed cropping, VLM-based landmark detection, DINOv3 embedding deduplication, and second-pass validation.These four stages translate empirical landmark-attention findings into a scalable street-view processing workflow.
  • Gaze-Informed Cropping: Panoramas are cropped to forward-facing regions based on good navigators’ gaze distributions, using 180° crops outside intersections and 270° crops at intersections.Intersection crops are widened to cover approaching and turning directions, while a narrower guideline region marks preferred detection areas.
  • VLM Detection: The VLM prioritizes distinctive, describable landmarks from qualitative categories, limits detections per batch, and returns structured bounding boxes, descriptions, categories, and navigational context.The prompt favors landmarks whose centers fall within the guideline region while permitting high-priority landmarks in the periphery.
  • Deduplication: DINOv3 embeddings compare detections across nearby frames and segment boundaries to reduce duplicate landmark predictions.The comparison uses cosine similarity of L2-normalized CLS-token embeddings within a ±5-frame window.
  • Pipeline Evaluation: VLM detections in some examples overlapped gaze hotspots, while other hotspots corresponded to visually attended features absent from verbal landmark descriptions.Examples of unmatched hotspots included store interiors, parked cars, and sidewalks.
  • Validation: Second-pass validation excludes detections that are invisible, mismatched, transient, insufficiently distinctive, or among generic infrastructure and temporary features.Validated detections must refer to the same visible object and be suitable as landmarks.
  • MR Interface: The interface presents landmarks as colored circles and uses directional arrows, gaze-triggered pulsing, and fading to guide attention during navigation.Users proceed to the next panorama after gazing at the highlighted landmark.

5 Preliminary User Evaluation

A within-subjects VR evaluation compared LandmarkLens with turn-by-turn navigation among eight poor-SOD participants. The study measured scene recognition, route choice, map sketches, and subjective responses using mixed statistical models and sketch-based spatial metrics.

  • Study Design: Eight participants with poor SOD completed routes under both LandmarkLens and turn-by-turn baseline conditions in a within-subjects study.All participants scored ≤3.5 on the SBSOD scale, and condition and route order were counterbalanced.
  • Study Design: The evaluation used new imagery from Ningyocho and two routes balanced for length, structure, and turn directions.Route A was 474.2 m and Route B was 410.9 m; both contained six segments and five turns.
  • Conditions: Both conditions provided joystick-based VR navigation and text turn instructions, while LandmarkLens additionally displayed VLM-generated gaze-interactive landmark highlights.Landmarks appeared as colored circles overlaid at predicted locations, accompanied by directional guidance.
  • Procedure: The procedure included introduction, practice, route exploration with tasks, and post-study debriefing.Each session lasted approximately 60 minutes and included scene recognition, route choice, and sketching-related instructions.
  • Measures and Analysis: Scene-recognition and route-choice performance were analyzed with mixed logistic regression, while map sketches were scored for route completeness, landmark count, and shape similarity.Shape similarity used bidimensional regression after translation, scaling, and rotation.
  • Measures and Analysis: Ordinal responses used mixed ordinal logistic regression, and open-ended responses were summarized through themes developed from open coding.The qualitative analysis focused on high-level themes from interviews and questionnaire responses.

6 Findings

LandmarkLens improved scene recognition and landmark recall, but did not significantly improve route-choice or structural map accuracy. Participants generally favored the prototype, while highlight volume and lack of personalization could interfere with route encoding.

  • Prototype use: 6m 24s versus 3m 55s: participants took longer with LandmarkLens than baseline because most dismissed the landmark circles before advancing.The gaze-to-dismiss interaction was optional, but most participants used it.
  • Scene recognition: 79.2% versus 64.9%: LandmarkLens significantly improved scene-recognition accuracy over baseline.The mixed logistic regression showed a significant condition effect, p < .01.
  • Route choice: 67.9% versus 73.2%: route-choice accuracy was comparable between LandmarkLens and baseline, with no significant condition effect.The reported significance test gave p = .535.
  • Subjective feedback: Participants reported that too many or non-personalized highlights could divert attention from routes and reduce recall of landmark-direction relationships.Participants with limited Japanese reading proficiency found unreadable signage difficult to retain.
  • Map sketching: 4.5 versus 2.5: participants recalled nearly twice as many landmarks in prototype map sketches as in baseline sketches.Landmark count differed significantly, p = .016, while route completeness and structural accuracy did not.
  • Subjective feedback: Six of seven post-route self-assessment items trended in favor of the prototype, although no individual item showed a statistically significant condition effect.Participants also rated the circles and directional arrows positively and generally found the system easy to learn.

7 Discussion

LandmarkLens partially supported spatial knowledge acquisition by improving scene recognition, while route and survey knowledge remained less affected. The study’s scope and real-world deployment conditions limit generalizability.

  • Guided Landmark Attention Partially Supports Spatial Knowledge Acquisition: LandmarkLens primarily detects visual landmarks, aligning with scene-level recognition but not the structural and semantic landmarks underlying route and survey knowledge.Participant feedback likewise described highlights as most useful at turns and overwhelming along straight segments.
  • Guided Landmark Attention Partially Supports Spatial Knowledge Acquisition: Repeated landmark-level gains could potentially scaffold route and survey knowledge over time, but this longitudinal effect remains an open question.The discussion identifies longitudinal investigation as a future direction.
  • Scope and Future Work: The final evaluation involved eight participants and routes within Tokyo, limiting evidence about generalizability across samples, regions, landmark densities, and urban layouts.The authors call for larger studies in other regions.
  • Scope and Future Work: Real-world deployment may introduce lighting variation, dynamic occlusion, and inference latency, despite offering richer spatial encoding through actual locomotion.These deployment conditions were not represented by the controlled VR evaluation.

8 Conclusion

LandmarkLens uses gaze-interactive VLM-generated overlays to direct people with poor sense of direction toward landmarks used by skilled navigators. It improved scene recognition and landmark recall, while route and survey knowledge remained comparable.

  • Findings: A 20-participant landmark attention study found broader scene scanning and more durable wayfinding cues among good navigators, while poor navigators referenced more transient features.These findings were translated into the system’s landmark prediction and overlay design.
  • Evaluation: LandmarkLens significantly improved scene-recognition accuracy and nearly doubled landmark recall in map sketches compared with a turn-by-turn baseline.Route and survey knowledge remained comparable between conditions.
  • Future direction: Participants valued highlights most at decision points and recommended semantically meaningful, personalized, context-sensitive highlighting.The conclusion connects these preferences to future adaptation by intersection context, familiarity, and individual preferences.

A Landmark Attention Study Materials & Additional Analysis

The landmark attention materials covered routes, participant groups, gaze dispersion, sketch-map comparisons, visual-element categories, and participant demographics and route presentation orders across Tokyo neighborhoods.

  • Route materials: Each neighborhood included one common route navigated by all 20 participants and ten unique routes assigned to individual participants.The route-selection materials are summarized in Figure 14.
  • Participant materials: Participant tables documented identifiers, demographics, SBSOD-based groups, language and Japanese-reading information, and route familiarity.Route positions R1–R8 were linked to specific routes and familiarity used a 1–10 scale.
  • Route materials: The presentation-order table listed ten counterbalanced orders, pairing one good and one poor navigator per order while alternating common-route presentation within neighborhood blocks.The route labels corresponded to Figure 14.
  • Gaze analysis: Gaze-dispersion frames plotted primary fixation locations for good and poor SOD participants using distinct teal and pink dots.The frames were selected for the highest within-group dispersion among good navigators.
  • Sketch-map analysis: Sketch-map examples juxtaposed actual route geometry with good- and poor-SOD participant sketches, with good navigators tending to preserve geometry and relative spatial relationships.The examples covered two common routes.
  • Verbal landmark analysis: The visual-element table paired category tags with example descriptions and reported percentages separately for good and poor navigators.These materials supported comparison of the groups’ articulated landmark types.

B VLM Pipeline Prompts

The VLM pipeline prompts prioritize permanent, distinctive, navigation-relevant landmarks while adapting selection to scene context and viewing direction. A second validation prompt checks whether detections correspond to recognizable, permanent objects before presentation.

  • B.1 Detection Prompt: The analyst identifies permanent, visually distinctive objects that help users orient themselves and remember routes.
  • B.1 Detection Prompt: Landmark priority is context-dependent: a lower-ranked object can outrank a higher-ranked one when it is the scene’s only distinctive feature.
  • B.1 Detection Prompt: Each image normally yields zero to two detections, with every image key represented and the same real-world object appearing at most twice across images.
  • B.1 Detection Prompt: The prompt favors objects near the walking direction and permits peripheral objects only when they are major, highly distinctive, or the only good candidates.
  • B.1 Detection Prompt: Bounding boxes target the most distinctive feature rather than the whole object, such as a mural, gate, logo, or station sign.
  • B.2 Validation Prompt: Validation combines category, description, and approximate bounding-box location, allowing nearby or partially overlapping objects when they refer to the same recognizable landmark.
  • B.2 Validation Prompt: Detections are rejected when objects are transient, generic, unrecognizable, mismatched, or represented by temporary signs and undistinctive building facades.

C User Evaluation Materials & Additional Analysis

The evaluation materials measure participants’ confidence, perceived spatial knowledge, route memorability, awareness, effort, and reactions to LandmarkLens highlighting. They combine route-choice questions with Likert-scale ratings of navigation and interface experience.

  • User Evaluation Materials: Participants rated confidence in scene recognition, intersection recall, mental-map accuracy, directional ability, environmental awareness, wayfinding without turn-by-turn directions, and mental effort.
  • User Evaluation Materials: Before ratings, participants selected which of two routes they would remember better for navigating again the next day.
  • User Evaluation Materials: Interface ratings covered highlight noticeability, distraction, arrow usefulness, dismissal naturalness, navigation experience, learning speed, and willingness to use AR glasses.

C.3 Participant Table

The participant table organizes demographics, SBSOD scores, and route familiarity, while the map-sketching results compare drawn route representations with actual route geometry using bidimensional correlation.

  • Participant Table: Participant records include PID, age, gender, SBSOD score, languages, Japanese fluency, reading ability, and familiarity with Routes 1 and 2.
  • Map Sketching Analysis: For each condition, map-sketching columns show route geometry, the original hand-drawn sketch, its vectorized form, and an overlay with bidimensional correlation coefficient r.
  • Map Sketching Analysis: Higher r values indicate greater spatial correspondence between the sketched route and the actual route geometry.
Loading 2608.30142v1…