Source-linked AI summary

Q&A or Document-Based? The Effects of Interface Type on How Screen Reader Users Access Interconnected Documents

Colleen F. Cipriano, Yichun Zhao, Miguel A. Nacenta, Kotaro Hara, Jaylee Soh

arXiv:2608.25382v1cs.HCcs.AIcs.IR

TL;DR

The paper asks how document-based and conversational LLM interfaces affect BLV screen reader users’ exploration and construction of interconnected knowledge. In a mixed-methods comparison, 16 users explored two fictional worlds with both interfaces. The DI supported broader document coverage, more accurate mental models and better knowledge application, although many participants preferred and perceived the QAI as better.

  • Problem

    It is unclear how LLM interfaces support or hinder BLV users’ ability to build interconnected knowledge from documents.

  • Method

    The study compared DI and QAI use by 16 screen reader users exploring two unfamiliar fictional worlds, using browsing logs, concept maps, decision tasks and interviews.

  • Results

    The DI produced wider document coverage, larger and more correct mental models, and better knowledge application, while many participants preferred the QAI despite weaker observed performance.

  • Takeaways & Limitations

    Interfaces that feel easier may support weaker understanding, so accessible information tools should be evaluated for accurate understanding and meaningful exploration as well as convenience.

  • Takeaways & Limitations

    The findings are limited by testing only screen reader users in two 90-minute sessions, and prolonged use may change interaction patterns and outcomes.

Abstract

from arXiv · show

Blind and low-vision (BLV) users are increasingly engaging with large language model (LLM) interfaces to access documents, but it is unclear how such systems support or hinder their ability to build interconnected knowledge. To examine this gap, we compared a Question-Answer Interface (QAI) that supports open-ended conversational inquiry, with a Document Interface (DI) based mostly on traditional structured text document navigation. We recruited 16 BLV screen reader users where they used both interfaces to explore two fictional worlds. Data from interaction logs, concept maps, decision-based tasks, and semi-structured interviews provide comparative insights into how interface design supports knowledge construction. Findings show that participants visited more distinct documents with the DI and formed larger and more correct mental models with the DI than with the QAI. They were also more able to apply knowledge they had gained. Simultaneously, many still preferred the QAI and often estimated that they had explored more, formed better mental models and applied their models better when acquiring the information with the QAI, despite this not being the case. Our analysis suggests possible interface design reasons for these differences and highlights some of the risks introduced by using question-answer interfaces to access information spaces.

1 Introduction

This paper compares document-based and conversational interfaces for screen reader users accessing interconnected documents. It examines how interface type affects browsing, knowledge integration and application, and interface preference.

  • The study compares a Document Interface (DI) with a Question and Answer Interface (QAI) for screen reader users exploring interconnected information.
  • The research asks how interface type changes browsing patterns, knowledge integration and application, and user preference.
  • 16 screen reader users explored two fictional worlds through both interfaces in a four-phase mixed-methods study.The phases measured browsing patterns, mental-model construction, knowledge application and interface reflections.
  • Participants covered a wider document range with the DI, while the QAI supported broader connections across fewer documents.
  • DI users produced larger, broader concept maps with fewer errors, whereas QAI maps were more densely linked.
  • Many participants preferred the DI’s explicit structure or the QAI’s conversational style, and subjective perceptions diverged from observed performance.

2 Background and Related Work

The background frames mental models and concept mapping as tools for studying document comprehension, while contrasting structured browsing with LLM-mediated conversational access. It identifies limited evidence about how LLM intermediaries shape accessible exploration of interlinked documents.

  • Mental Models: A mental model is an internal representation that organizes knowledge about the world and supports action based on that knowledge.
  • Mental Model Elicitation: Concept mapping externalizes how people organize concepts and relationships, allowing comprehension to be examined beyond simple memory retrieval.
  • Information Access: Traditional information retrieval distinguishes search, which returns results for a query, from browsing across document collections.
  • Conversational Access: GenAI systems can interpret nuanced input, synthesize text across documents and maintain conversational context during information exploration.
  • Research Gap: Little is known about LLMs as intermediaries between users and documents, especially for accessibility-centered exploration of interlinked documents.
  • Risks: Prompt-and-summary interaction may reduce exposure to wider document collections, with prior work associating it with weaker exploration, comparison and continuity.
  • Research Motivation: The paper therefore compares document-based and question-answer interfaces while aiming to make documents and interfaces as accessible as possible.

3 Study Methodology

The study used unfamiliar fictional worlds and a within-subjects comparison of document-based and question-answer access. Participants explored multimodal document sets containing textual articles, links, diagrams and tactile spatial representations.

  • The study observed screen reader users navigating unfamiliar information spaces, constructing mental models and applying acquired knowledge across two laboratory sessions.
  • Study Materials: Researchers created Solana and Dominion as fictional, interconnected document worlds to minimize prior-knowledge effects and support extended exploration.
  • Study Materials: Each world contained 25 documents, including an index, Wikipedia-like textual articles, meaningful links and diagrams about history, governance and culture.
  • Study Materials: Spatial documents were made accessible with 3D-printed tactile overlays designed according to tactile-representation best practices and accessibility guidelines.
  • Study Procedure: The procedure included interface conditions, verbal elicitation, concept mapping, a decision-based task and interface reflection.
  • Study Procedure: Participants completed all phases for both DI and QAI conditions in counterbalanced order.
  • Study Setup: The DI setup combined a tactile overlay on a tablet with a laptop displaying the Dominion interface.
  • Study Materials: The fictional-world documents and tactile versions were open sourced for subsequent research.

3.2 Interfaces (Main Condition)

The DI provided direct exploration of accessible hyperlinked documents and tactile materials, whereas the QAI mediated corpus access through conversational question answering. The QAI constrained responses to the document knowledge base and offered several output formats.

  • The DI and QAI were designed as state-of-the-art alternatives representing direct corpus exploration and LLM-mediated question answering.
  • Document Interface: The DI used WCAG 2.1 web pages with index and in-page search through participants’ own devices and screen readers.
  • Document Interface: Spatial documents in the DI used 3D-printed overlays whose touched areas generated verbal descriptions alongside tactile relief.
  • Document Interface: Tactile representations preserved spatial relationships that are difficult to communicate through screen-reader navigation alone.
  • Question-Answer Interface: The QAI was a WCAG 2.1 web chatbot accepting typed or spoken questions and returning text, bullets, tables, image descriptions and quotations.
  • Question-Answer Interface: The conversational system loaded 25 corpus documents into its context and used cache-augmented generation to constrain answers to the knowledge base.
  • Question-Answer Interface: When answers were unavailable, the QAI suggested related documents; for out-of-scope questions, it stated that information was unavailable and offered alternatives.
  • Question-Answer Interface: The study reported no observed hallucinations or erroneous responses from the QAI.

3.3 Participants

The study recruited 16 regular screen reader users and had them explore two fictional worlds through both a document interface and a question-answer interface. Across two sessions, participants explored openly, constructed researcher-assisted concept maps, applied their understanding to scenarios, and discussed their experiences.

  • Participants: 16 participants reported regular use of at least one screen reader, including JAWS, NVDA, VoiceOver, or TalkBack.They were recruited through community organizations, social media, and snowball sampling.
  • Procedure: Participants completed two 90-minute sessions on separate days, with interface and world order counterbalanced.Sessions were separated to reduce cognitive load, fatigue, and carryover effects.
  • Procedure: During exploration, participants followed their interests and prioritized acquiring domain knowledge over constrained recall performance.Before QAI sessions, they were told that the system contained 25 documents.
  • Procedure: Participants described each world, identified important concepts and relationships, and validated researcher-constructed concept maps without access to the assigned interface.The researcher used probing questions and a structured script to refine the maps.
  • Procedure: Participants applied their understanding to world-specific decision scenarios and then completed interviews about ease of use, challenges, preferences, exploration, learning, and confidence.The interviews included open-ended explanations of their comparative judgments.

3.5 Measurements and Analysis

The study combined interaction-log analysis, concept mapping, decision-based assessment, graph measures, error coding, and thematic analysis to compare how participants explored and represented the fictional document worlds.

  • Interaction Logs: Phase 1 logged visited documents and transitions, measuring unique document visits, total document visits, unique transitions, total transitions, and transitions per document.These measures quantified participants’ navigation patterns.
  • Interaction Logs: Visits were operationalized differently across interfaces: opening a document counted in DI, while QAI responses were mapped to source documents.A QAI source counted as visited when its information covered five or more contiguous sentences.
  • Graph Analysis: Exploration graphs represented topics as nodes and transitions as edges, with density, directed diameter, degree of variance, and average shortest path length as structural measures.These metrics captured interconnectedness, breadth, connectivity evenness, and path length.
  • Concept Maps: Concept maps recorded topics, subtopics, and explicit connections, counting only unambiguous references mapped to predefined documents and marking beyond-content references as inaccurate.Repeated mentions did not create new instances unless they introduced new connections.
  • Concept Maps: Error metrics counted contradictions, unsupported inferences, and incorrect detail recall, excluding repeated errors unless they added a distinct claim or connection.Graph metrics were also calculated for participants’ mental models.
  • Decision Tasks and Qualitative Analysis: Decision-based tasks used topic, subtopic, connection, and error measures but omitted graph metrics because responses were sparser and some participants could not produce them.Qualitative analyses used logs, queries, responses, visualizations, and interviews, with thematic coding for later research questions.

3.6 Positionality Statement

The authors state that the study centers BLV screen reader users’ lived realities while acknowledging that none of the authors identify as BLV.

  • Positionality: None of the authors identify as BLV, but the study grounds its design, analysis, and interpretation in participants’ accounts, practices, and perspectives.The authors state that one author has more than six years of sustained engagement with BLV communities.

4 Results

The Results section organizes study evidence by the research questions and presents quantitative measurements before qualitative analysis.

  • Results: Evidence is organized by research question, with quantitative measurements considered before qualitative analysis within each subsection.

4.1 How do People Explore Information Spaces with the Different Interfaces? (RQ1)

The DI supported broader document coverage, while the QAI created tighter connections across topics within a smaller document set. Qualitative findings indicate that each interface fit different exploration goals depending on whether participants accepted the documents’ predefined structure or wanted to define their own.

  • Quantitative Results: 22.9% more distinct documents were visited with the DI than the QAI, despite similar total visit counts.Participants visited an average of 11.44 documents with the DI and 9.31 with the QAI.
  • Quantitative Results: The QAI produced denser navigation graphs than the DI, indicating tighter connections between different topics.The reported mean graph density was 0.21 for the QAI versus 0.13 for the DI.
  • Qualitative Results: Participants used overview, depth, and fact finding as recurring exploration patterns across the two interfaces.The qualitative analysis identified these three patterns as recurring forms of engagement with the information spaces.
  • Qualitative Results: The DI helped participants who accepted a predefined structure, while the QAI helped those seeking a self-defined organization of information.The comparison applied across overview and depth, with interface effectiveness depending on participants’ preferred information organization.
  • Qualitative Results: The QAI readily supported fact finding and cross-document summaries, whereas the DI made participants hunt through multiple documents for specific answers.Participants described the QAI as providing contextualized facts and organizing information around their interests without checking every document.
  • Qualitative Results: The DI offered broader coverage and an explicit structure, while the QAI supported coherent answers assembled from information scattered across documents.Participants struggled to obtain an overview with the QAI when they did not know the information space well, but could use it for more customized exploration.

4.2 How do the Different Interfaces Influence Integration of Knowledge? (RQ2)

The DI supported broader exploration and larger, more accurate mental models than the QAI, while the QAI helped participants connect related information more densely. These measured advantages extended to applying knowledge, although participants faced different interface-specific challenges.

  • Concept Maps (Phase 2): 67% more topics were included in DI mental models than QAI models, while DI also produced 11% more subtopics overall.The combined topic and subtopic counts indicate more comprehensive models with DI; the subtopic difference was not statistically distinguishable.
  • Concept Maps (Phase 2): DI mental model graphs were less interconnected than QAI graphs, with lower density and higher diameter, average shortest-path length, and degree variance.These structural differences were calculated after removing errors.
  • Decision-Based Application (Phase 3): The scenario task showed a similar DI advantage, but six participants could not provide coherent responses and higher DI connections were not statistically distinguishable.Graph metrics were omitted because the resulting scenario graphs were sparse and considered noisier and less meaningful.
  • Qualitative Results: Participants associated DI with stable structure and visual organization, whereas QAI supported self-directed integration but required skill in formulating effective questions.DI users described more information to work with, while QAI users valued coherent answers spanning related topics but sometimes questioned their completeness.

4.3 What Aspects of the Different Interfaces Affect Preferences? (RQ3)

Participants’ preferences reflected a trade-off between DI’s explicit control and lower guesswork and QAI’s conversational flexibility. These perceptions often diverged from measured exploration, learning, and confidence outcomes.

  • Comparative Reflections: Perceived exploration, learning support, and confidence did not fully align with measured performance, with the discrepancy strongest for learning and answering confidence.Participants’ subjective impressions and preferences therefore did not consistently track the observed interface differences.
  • Comparative Reflections: Preferences divided between DI’s explicit structure and QAI’s conversational style, despite measured advantages for DI in breadth and mental-model quality.The interviews identified agency/control and effort/cognitive load as reasons supporting preferences for either interface.
  • Agency and Control: Participants valued DI for choosing documents and links directly, while QAI control depended on knowing how to steer the conversation productively.Nine participants valued DI’s navigation control, while six described conversational steering as important to control.
  • Effort and Cognitive Load: DI was often viewed as straightforward and low-guesswork, whereas its effort could involve actively processing and organizing information.Participants used effort and cognitive-load considerations to support preferences for either interface.

5 Discussion

The discussion identifies a trade-off: DI encouraged broader document coverage and stronger mental models, while QAI offered flexible, efficient-feeling interaction that often obscured weaker learning outcomes. It recommends user awareness and hybrid designs, while qualifying conclusions through measurement and study-design limitations.

  • RQ1: DI led participants to visit more unique documents, whereas QAI enabled more flexible transitions across the information space.The authors relate this trade-off to the interfaces’ structural differences.
  • RQ2: QAI produced smaller, more error-prone mental models and did not improve participants’ ability to apply acquired knowledge.The authors note that QAI models may have been more tightly connected despite being poorer overall.
  • Interpretation: The authors propose that DI’s greater effort may encourage elaboration and deeper processing, while QAI’s removal of predetermined structure may weaken learning.These are presented as plausible causal explanations rather than established mechanisms.
  • RQ3: Participants often rated QAI highly for outcomes where their measured performance was worse, and preferences were evenly split between the interfaces.This pattern suggests that interface characteristics can obscure value to users or shape preference through factors beyond measured outcomes.
  • Implications: Users should recognize that QAI can feel efficient and flexible while increasing risks of superficial processing and weaker knowledge application.The authors emphasize this risk for BLV users seeking to reduce the burden of traditional document access.
  • Implications: Future systems could combine document navigation, search, and conversational interaction to support different tasks and preferences.The study separated DI and QAI to make their differences easier to identify, whereas hybrid systems may combine their advantages.
  • Limitations: The DI advantage cannot be attributed to interface type alone because multimodal document access and corpus structure may also have influenced performance.The authors specifically identify touch interaction for spatial documents and an index document linking the corpus as potentially influential design choices.
  • Limitations: Phase 1 findings depend partly on pre-analysis choices about counting visits and transitions, although later model-elicitation and scenario measures were less variable.The authors state that alternative counting choices might have changed some Phase 1 balances.

6 Conclusion

The study compares DI and QAI for BLV users exploring unfamiliar domains and finds a consequential mismatch between perceived and measured benefits. QAI can feel expansive and easy, while DI’s structure supports more comprehensive mental models, making evaluation beyond convenience important.

  • Conclusion: QAI’s conversational interaction can feel deceptively expansive, but this did not translate into better learning.DI may demand more effort or feel overwhelming while producing a more comprehensive mental model.
  • Conclusion: Many participants preferred QAI even when performance suggested weaker understanding, creating a risk that easier-feeling interfaces support less accurate knowledge.The authors recommend evaluating accessible technologies for accurate understanding and meaningful exploration, not only speed or convenience.
Loading 2608.25382v1…