Source-linked AI summary
MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places
Jason Armitage, Ioannis Tsochantaridis, Linda Mazzone, Chuqiao Yan, Srini Narayanan, Sarah Ebling
TL;DR
People with accessibility requirements often need incomplete or outdated information when planning visits to real-world places, while the safety and usefulness of multimodal AI assistance remain unclear. MAP introduces a benchmark with claim verification and visual evidence retrieval, comparing system responses with refreshed ground truth and human checks. Results show that search access can increase “Useful and Safe” responses, while visual-evidence performance varies across systems and accessibility features.
Problem
Accessibility information for planning real-world activities can be incomplete, distributed across sources, or outdated, and the safety and usefulness of multimodal AI assistance remain unclear.
Method
MAP benchmarks multimodal AI assistants on claim verification and visual evidence retrieval for accessibility planning, using a time-based ground-truth workflow with automated construction and human reviews.
Results
Search access increased GPT-5.2 outputs labeled “Useful and Safe” from 17 to 141, while visual-evidence performance varied across systems and features.
Takeaways & Limitations
MAP evaluates factual accessibility information and relevant visual evidence in an open-world setting where places and accessibility conditions can change over time.
Takeaways & Limitations
The benchmark relies on public digital sources whose coverage favors visible physical and mobility features and may lag behind changes in the physical world.
Abstract
from arXiv · showhide
We introduce MAP, the first benchmark to evaluate multimodal AI systems as assistants for users with accessibility requirements when planning visits to places in the real world. In our evaluation, systems are presented with requests to verify or recommend a point of interest meeting an accessibility requirement. MAP contains two novel assessments: Claim verification for accessibility planning assesses if information on places and stated accessibility features is supported and identifies places that satisfy requested accessibility features. Visual evidence retrieval for accessibility planning checks if a multimodal AI system can select visual evidence for the requested place and accessibility feature. Our methodology supports comparison of AI systems in a setting where place information and accessibility information can change over time by evaluating systems and refreshing ground truth data at scheduled times. The benchmark is based on automatic rating and human rating for a proportion of responses.
1 INTRODUCTION
MAP addresses the challenge of helping people with accessibility requirements plan visits in a complex, changing urban environment. It benchmarks multimodal AI systems on accessibility information, evidence, and place recommendations against refreshed ground truth.
- Accessibility information needed for planning may be incomplete, distributed across sources, or outdated.Examples include step-free entrances, tactile guidance, audible signals, lighting conditions, and visual information.
- MAP evaluates multimodal AI systems as assistants for users with accessibility requirements when finding real-world places to visit.Responses are compared with ground-truth data on accessibility features, using large multimodal AI systems and human checks.
- Refreshed ground truth data and fixed scoring procedures support comparisons between systems as place and accessibility information changes over time.Relevant changes include updates to physical access, new or closed places, and revised evidence about accessibility features.
- The benchmark includes claim verification and visual evidence retrieval for accessibility planning.Claim verification assesses correctness and justification of accessibility-related responses, while visual evidence retrieval assesses whether relevant supporting images are found.
- MAP covers accessibility requirements associated with mobility, vision, hearing, and cognition.The benchmark’s reference data maps evidence to a single accessibility schema across these requirements.
2 BACKGROUND
Prior multimodal and accessibility evaluations provide useful evidence but often use static settings or focus on one disability type and application domain. MAP extends this work with a broader benchmark for up-to-date accessibility information about real-world places.
- Prior multimodal benchmarks evaluate tasks such as multimodal understanding, factuality, recommendation, verification, and navigation.Examples include MMMU, MMBench, LiveBench, CitySeeker, V-IRL, NavBench, AgentRecBench, AgentSelect, FACTS, RW-post, and MUSCICLAIMS.
- Existing spatial and navigation benchmarks commonly use static scenarios with a pre-selected place or point of interest and fixed ground-truth information.MAP addresses this limitation by assessing up-to-date and factually correct accessibility information in an urban centre.
- Prior accessibility evaluations include built-environment audits, virtual navigation for blind users, micro-navigation queries, and hearing-related chatbot assessments.These studies span different application domains but do not provide the same broader real-world planning framework described for MAP.
- Existing accessibility studies are constrained to one disability type and a specific application domain.MAP extends beyond this limitation with multiple disability types and updated accessibility information for real-world locations and diverse features.
- Automatic evaluation metrics in related multimodal tasks are often complemented by human evaluation.Recommendation work commonly uses top-N accuracy, while verification work commonly uses accuracy and F1 score.
3 THE MAP BENCHMARK
MAP benchmarks multimodal AI assistants for accessibility-aware visit planning through claim verification and visual evidence retrieval. It uses structured prompts, a unified accessibility schema, multimodal ground truth, and refreshed data for real-world place information.
- MAP evaluates whether multimodal AI systems provide accurate information about places and accessibility features for users planning real-world visits.
- The benchmark contains claim verification for feature checks and search-by-feature recommendations, plus visual evidence retrieval from candidate images.
- Four disability categories are covered: mobility, vision, hearing, and cognition, with prompts specifying accessibility features and either a place or place type.
- The first benchmark version focuses on Zurich and includes 59 accessibility features across 257 place types in 34 local areas.
- MAP constructs POI ground truth by integrating multiple sources into one schema, mapping textual and visual evidence to fields, and resolving conflicts using relevance, recency, source type, and agreement.
- Ground-truth context is assembled by prompt type, while visual retrieval uses retained images with descriptions and feature-visibility information.
4 EVALUATION
MAP evaluates accessibility planning through claim verification and visual evidence retrieval, using ground-truth comparisons, automatic rating, and human checks. Results show cautious or unresolved responses remain common, while search improves useful-and-safe recommendations for several closed-weight systems but can also increase potentially unsafe outputs.
- Claim verification: Claim verification checks whether named places and accessibility claims match ground-truth evidence, while search-by-feature additionally requires identifying suitable places.Feature-check prompts assess one named POI; search-by-feature prompts require finding POIs that satisfy access requirements.
- Claim verification: Automatic rating applies policy-defined checks against consolidated ground truth, with human rating used to externally check a proportion of responses.The evaluation uses label-based scores describing usefulness and whether accessibility information is safe to rely on.
- Claim verification: Across scored feature-check responses, “Unverified” is dominant, indicating that systems often avoid definite accessibility claims even when evidence is retrieved.Open-weight systems remain mostly “Unverified” in both search conditions, with retrieved evidence producing cautious responses more often than “Safe” ones.
- Claim verification: Search increases “Useful and Safe” outputs for several closed-weight systems, including GPT-5.2 rising from 17 to 141 responses.Search also increases “Useful but Potentially Unsafe” outputs, suggesting more specific and decisive responses alongside the gains in verified utility.
- Rating checks: Human and automatic ratings align for several systems, but the widest label-allocation gaps occur for Grok 4.3 and Grok 4.5.Evaluation challenges include extra features, contradictory information, and insufficient specificity in responses.
- Visual Evidence Retrieval for Accessibility Planning: Visual evidence retrieval asks systems to select an image or None for a POI and accessibility feature, scoring place correctness and visibility of the relevant area.The evaluation compares selections with ground-truth candidate metadata and assigns a visibility score from 0 to 1.
- Visual Evidence Retrieval for Accessibility Planning: Gemini 3.6, Gemini 3.1 Pro Preview, GPT-5.6 Terra, and Claude Sonnet 4.6 each select correct visual evidence in more than two thirds of cases.Other systems either select None or choose distractors showing the requested feature area at the wrong POI; performance also varies by feature.
5 LIMITATIONS
MAP’s limitations concern its digital ground truth, prompt design and feature coverage, and Zurich-only evaluation scope. These constraints affect how broadly and currently its results can be interpreted.
- Ground truth: Digital ground truth favors visible physical-access features over hearing- and cognition-related features because those features appear more often in online sources.The ground truth is built from publicly available digital sources, producing uneven accessibility-feature coverage.
- Ground truth: Digital sources may lag behind physical-world changes, although MAP refreshes ground truth before scheduled evaluation runs.Renovations, temporary works, closures, and visitor-support changes can make websites, reviews, images, and open data temporally inconsistent.
- Evaluation scope: Figure 4 reports correct visual evidence rates for matching a requested accessibility feature in a point of interest, with full label counts in Table 4.The figure concerns the test partition of visual evidence retrieval.
- Prompt design: Prompts may use synthetic American English and may not represent shorter, underspecified, error-prone, regional, cultural, or multilingual user requests.This limitation reflects choices made for the first MAP iteration.
- Prompt design: MAP prompts are limited to schema features that can be assessed and scored, excluding some temporary constraints and journey-level access requirements.Real requests may concern features outside the benchmark’s coverage or accessibility along a route rather than at a place.
- Evaluation scope: Both prompts and ground truth are limited to Zurich, so results should be interpreted as evidence for that city rather than accessibility planning across regions.The geographic scope is restricted in this first iteration.
6 CONCLUSION
MAP evaluates multimodal AI assistants for accessibility planning in real places through factual and visual-evidence tasks. It supports open-world evaluation by refreshing ground truth as place information changes.
- Conclusion: MAP evaluates whether multimodal AI systems provide factually accurate accessibility information for visits to real places.The evaluation compares responses with ground truth data and includes automatic and human assessment.
- Conclusion: The benchmark evaluates visual evidence relevant to a user’s accessibility request alongside factual accuracy.Its two tasks assess information correctness and whether presented images match the requested accessibility feature.
- Conclusion: MAP updates ground truth data and runs assessments within specified windows to accommodate dynamic changes in the real world.This enables evaluation in an open-world setting rather than relying on permanently fixed place information.
1 QUESTION TYPES AND SYSTEM PROMPTS
The benchmark uses sampled question prompts served as user requests and separate standardized system prompts. Tables 5 and 6 document these prompt materials by question type.
- Question types: Samples of prompts for each question type are presented in Table 5 as lists of elements constituting the task prompts.The prompts are designed to align inference with how human users query systems.
- System prompts: At inference, question prompts are served as user requests, while a separate system prompt instructs models to produce outputs suitable for evaluation.The system prompt is distinct from the user-facing request.
- System prompts: System prompts are concise and standardized across evaluated models, with their configurations listed by question type in Table 6.Table 6 records the system prompts used for the reported runs.
2 ADDITIONAL RESULTS
Additional results report validation-partition runs across the three question types spanning MAP’s claim verification and visual evidence retrieval tasks.
- Validation results: Validation-partition results are reported for all three question types.The section organizes these results across the benchmark’s question-type structure.
- Validation results: The reported validation runs cover both claim verification and visual evidence retrieval.These are the two evaluation tasks named in the passage.
- Validation results: Tables 7, 8, and 9 present the validation-partition results for the three question types.The results are distributed across three tables.
3 HUMAN ASSESSMENTS
Human assessment highlights how additional accessibility claims can change response labels, especially when they conflict with ground truth or cannot be verified. These cases matter because incorrect accessibility information can be consequential for users.
- Case-based scoring: Correctly answering the requested feature does not ensure a Safe label when an additional claim contradicts ground truth.In Case 1, correct wheelchair-accessible seating was overridden by a false entrance claim.
- Case-based scoring: Broad statements such as “wheelchair accessible” may be insufficient when the prompt requests a specific feature such as seating.The phrase can refer to an entrance, restroom, or other element rather than the requested seating.
- Validation: Validation results are reported separately for feature-check questions and search-by-feature prompts, using counts of auto-rater labels.The feature-check validation uses n=300 questions, while search-by-feature validation uses n=200 prompts.
- Case-based scoring: A response can become Unverified when it includes a feature claim absent from the ground-truth data, even if its core answer is correct.Case 3 includes correct restroom information but also an unverifiable parking claim.
- Assessment significance: Human assessment treats incorrect accessibility information as consequential for users planning visits.
4 SPECIFICATIONS FOR EVALUATION RUNS
The evaluation runs use independent, controlled inference and automatic scoring procedures across claim verification and visual evidence retrieval. Human-assessment examples illustrate how requested and additional accessibility claims are judged against ground truth.
- Evaluation settings: Each prompt is processed independently with stateless execution and no context from earlier prompts.
- Visual evidence retrieval: Visual evidence retrieval uses deterministic scoring based on one correct sample and controlled distractor images.
- Inference settings: Models generally run at temperature 0 with low reasoning where supported, while open-weight models are served through Ollama with specified token and hardware settings.Qwen2.5-VL-32B-Instruct uses num_predict=500; runs use a single NVIDIA L4 GPU with 24 GB VRAM.
- Automatic rating: Claim verification is automatically scored by extracting accessibility claims, mapping them to MAP fields, and comparing them with consolidated ground-truth values and evidence.The assessment considers accessibility features added beyond those requested in the prompt.
- Human-assessment examples: In one human-assessment case, a correct seating claim becomes Contradictory because an added wheelchair-accessible entrance claim conflicts with ground truth.
- Human-assessment examples: Other cases show that broad accessibility wording may miss the requested feature, while correct answers can remain acceptable when no wrong information is added.The examples concern wheelchair-accessible seating and restroom claims, including responses about Zurich locations.