Source-linked AI summary
Designing and Evaluating AI Margin Notes in Document Reader Software
Nikhita Joshi, Daniel Vogel
TL;DR
Document readers commonly isolate LLM capabilities in disconnected chat interfaces, leaving integrated AI commenting underexplored. Across three experiments, the paper evaluates AI margin-note integration, text-selection automation, and human–AI involvement, finding preferences for integrated notes and manual creation while involvement minimally affected comprehension.
Problem
LLM capabilities in document readers are usually disconnected from document text, leaving the design and evaluation of integrated AI comments underexplored.
Method
Three experiments evaluate AI margin notes across integration, text-selection automation, and varying degrees of human and AI involvement.
Results
Participants preferred integrated AI margin notes and manual selection, while AI involvement produced only marginal differences in reading comprehension.
Takeaways & Limitations
Document reader software should offer multiple AI margin-note variations to support different user goals and preferences.
Takeaways & Limitations
The findings, especially for reading comprehension, may not hold after extended long-term use or across broader educational settings.
Abstract
from arXiv · showhide
AI capabilities for document reader software are usually presented in separate chat interfaces. We explore integrating AI into document comments, a concept we formalize as AI margin notes. Three design parameters characterize this approach: margin notes are integrated with the text while chat interfaces are not; selecting text for a margin note can be automated through AI or manual; and the generation of a margin note can involve AI to various degrees. Two experiments investigate integration and selection automation, with results showing participants prefer integrated AI margin notes and manual selection. A third experiment explores human and AI involvement through six alternative techniques. Techniques with less AI involvement resulted in more psychological ownership, but faster and less effortful designs were generally preferred. Surprisingly, the degree of AI involvement had no measurable effect on reading comprehension. Our work shows that AI margin notes are desirable and contributes implications for their design.
1 Introduction
This paper introduces AI margin notes, integrating LLM-supported document comments with text, and evaluates their integration, text-selection, and human–AI involvement through three experiments. Participants preferred integrated margin notes and manual selection, while greater AI involvement improved speed and effort but reduced psychological ownership without affecting comprehension.
- 1 Introduction: AI margin notes extend established integrated note-taking practices, including marginalia and digital comments attached to selected document text.Readers traditionally write notes either separately from the document or in spaces integrated with its text; digital systems commonly position comments beside selected text.
- 1 Introduction: The paper explores AI margin notes as LLM-enhanced comments integrated with document text rather than delivered through a separate, disconnected chat interface.LLMs already support summarizing, rewording, and explaining unfamiliar concepts, but these capabilities had not been integrated into document-reader commenting features.
- 1 Introduction: Three controlled experiments examined three design parameters: integration with the document, automated versus manual text selection, and the degree of human and AI involvement in note generation.Participants read short nonfiction documents, interacted with an LLM primarily through comments, and answered comprehension questions two hours later.
- 1 Introduction: AI margin notes were preferred over chat interfaces, manual text selection was preferred, and AI involvement did not measurably affect reading comprehension.Techniques with more human involvement were associated with greater psychological ownership, whereas more AI involvement made techniques faster, less effortful, and generally preferred.
- 1 Introduction: The work contributes the concept of AI margin notes and design guidance for balancing integration, selection control, psychological ownership, speed, effort, and preference.The findings show that different human–AI involvement levels were valued for different reasons rather than producing a single universally superior technique.
2 Background and Related Work
Prior research shows that note-taking can support memory and learning but also imposes substantial cognitive and temporal demands. LLM-based reading tools may improve comprehension and reduce effort, yet AI margin notes remain underexplored and could either hinder or support deeper engagement with text.
- Note-taking benefits: Note-taking supports learning through both external information storage and the encoding benefits of processing and organizing ideas.Deeper processing, such as connecting text to prior knowledge and experiences, strengthens the encoding function.
- Note-taking challenges: Note-taking requires readers to coordinate reading and writing, creating mental and temporal demands that can delay reading and hinder comprehension.These demands can occur even without time limits and may become especially tiring during longer reading tasks.
- LLM-assisted reading: LLM-based reading tools have been used for simplification, skimming, question answering, summarization, and document note-taking, with some approaches improving comprehension or reducing workload [3] [27] [26].Examples include ChatGPT, Adobe Acrobat, Google Notebook LM, NoteGPT, and ChatPDF.
- LLM-assisted learning: Prior work suggests a tradeoff between learning and effort: independent and chatbot-assisted note-taking outperformed asking the chatbot questions alone, although the latter was preferred as more enjoyable and less effortful.Other LLM systems can personalize learning and improve academic performance, mental effort, or motivation, but learners may over-rely on them.
- Research gap and motivation: No prior work had thoroughly examined integrating LLMs into document-reader commenting, motivating study of AI margin notes as tools that might either hinder deeper processing or reduce note-taking demands and support engagement.Existing prototype mockups used selected text, predetermined prompts, and responses shown as sticky notes or beneath the selected text.
3 AI Margin Notes
AI margin notes are framed by three design parameters: integration with document text, selection automation, and the balance of human and AI involvement. Compared with separate chat interfaces, integrated notes can reduce prompting, attention-switching, and response-retrieval work [44].
- Design parameters: AI margin notes are characterized by integration, selection automation, and the level of human and AI involvement.Comments are anchored to specific text, unlike chat tools placed in a separate side panel disconnected from document context.
- Integration: Integrated notes can reduce the work of referring to specific passages, switching between reading and writing interfaces, and retrieving prior responses [44].Chat interfaces may require carefully formulated references or copy-paste steps, impose switching costs, and bury explanations in chat history.
- Human and AI involvement: AI margin-note text can range from entirely human-written to entirely AI-generated, with intermediate techniques combining human and LLM involvement.Greater human involvement may increase cognitive engagement, support reading comprehension, discourage verbatim notes, and improve psychological ownership [34] [35], although excessive engagement can frustrate and discourage readers.
- Selection automation: Selection for an AI margin note can be automated by the assistant or performed manually by the user, with automation reducing effort but potentially lowering psychological ownership and comprehension.Psychological ownership can support active and motivated learning through greater control over learning experiences.
4 Experimental Method
Three within-subjects experiments evaluated how AI margin-note integration, selection automation, and human–AI involvement affect comprehension, duration, ownership, workload, and preferences. Participants read approximately 500-word documents with an LLM technique, completed comprehension tests, and rated their experiences and preferences.
- Experimental Design: Three experiments used the same method with different participants to examine integration, selection automation, and human–AI involvement across comprehension, duration, ownership, workload, and preferences.All experiments were conducted through Prolific and used the same experimental method.
- Procedure: Participants first read approximately 500-word non-fiction documents while interacting with an LLM technique, then answered six multiple-choice comprehension questions.The reading and testing stages were separated, with the test stage occurring two hours later and using a 60-second closed-book limit.
- Procedure: Participants completed three margin notes or AI responses per document, rated each experience, ranked conditions, and provided overall preferences after trying all techniques.Overall condition rankings were derived using the Condorcet voting method, and additional experiment-specific behavioral metrics were introduced with the results.
- Experimental Design: All experiments used within-subjects designs with randomly assigned conditions and documents drawn from six documents by Wallace et al..Documents could be assigned to any condition to increase internal validity, allowing participants to compare techniques directly.
- Measures: Reading comprehension and duration came from interaction logs, while ownership, workload, and other measures came from post-reading questionnaires.Comprehension was scored as the number of correctly answered questions on a 0–6 scale; other metrics generally used 0–100 interval scales.
5 Experiment 1: Integration
Experiment 1 compared traditional chat-based prompting with AI margin notes integrated into the document. Although conditions did not differ significantly on comprehension, duration, ownership, workload, or use frequency, participants generally preferred the integrated notes because they were easier to reference and ask about.
- Experiment 1: Integration: The experiment compared a right-side chat interface with margin notes created by selecting document text, opening an anchored comment, and prompting an LLM.Participants could edit or delete comments, and clicking a comment highlighted its corresponding document text.
- Experiment 1: Integration: Participants preferred AI margin notes over chat: note ranked first overall, with 17 participants (65%) assigning it first place versus chat’s second-place ranking from 16 participants (62%).The Condorcet overall rankings were 1 for note and 2 for chat, despite no significant condition effects on the measured performance, workload, ownership, or usage outcomes.
- Experiment 1: Integration: Participants found margin notes easier and more intuitive because they avoided switching interfaces, embedded comments beside document text, and supported selecting text related to a question.Twelve participants (46%) specifically mentioned these advantages in free-form responses.
- Experiment 1: Integration: Note prompts used deictic references to selected text more often than chat prompts, with 29 prompts (37%) for note versus 17 (22%) for chat.Chat also contained 10 prompts (13%) whose deictic references pointed to nouns from earlier prompts, whereas note selections supplied more immediate textual specificity.
6 Experiment 2: Selection Automation … 6.3 Results
Experiment 2 compared AI-selected and manually selected text for generating three AI margin-note summaries. Manual selection increased time and effort but also psychological ownership, perceived performance, preference, and control, while reading comprehension did not differ.
- 6 Experiment 2: Selection Automation: Both conditions represented document summarization tasks normally performed by copying text into a chat interface or by pressing an automated button.The experiment isolated how selection automation affects the association between AI margin notes and specific text selections.
- 6.1 Participants: The experiment recruited 32 new Prolific participants, with 30 valid responses after excluding two participants for invalid or unsupported reading-comprehension behavior.Participants met Experiment 1’s inclusion criteria and received $10 plus a $5 bonus incentive for top-25% comprehension performance.
- 6.2 Apparatus: The automatic technique generated three summary comments at once through a toolbar button, while the manual technique required participants to select text and create three comments individually.Manual comments could also be deleted or regenerated individually.
- 6.3 Results: Reading comprehension showed no significant difference between manual and automatic selection conditions.The analysis therefore focused on duration, ownership, performance, effort, preference, and selection behavior.
- 6.3 Results: Manual selection was slower and more effortful than automatic selection, but participants preferred it and reported greater psychological ownership and perceived performance.Manual duration was 3.43 versus 2.78 for automatic, while manual psychological ownership was 50.25 versus 6.75 and perceived performance was 88 versus 80.5.
- 6.3.3 Task Workload: Manual selection increased perceived workload specifically on Effort, while Mental Demand, Physical Demand, Temporal Demand, and Frustration did not differ between conditions.Manual Effort had a median of 35 versus 14 for automatic.
- 6.3 Results: Manual selection was preferred for repeated use: its Frequency of Use median was 79.5 versus 55 for automatic, and 77% ranked manual first.Participants also tended to rank automatic second, with 70% assigning it a ranking of 2.
- 6.3.4 Preferences: Manual selection produced shorter selections distributed throughout the document, whereas the LLM’s automatic selections clustered near the beginning.Manual selections had a median of 62 words versus 83 for automatic selections.
6.4 Summary
Manual text selection for AI margin notes was slower and more effortful but associated with greater psychological ownership, better perceived performance, and participant preference. Although the resulting notes were always summaries, other human–AI techniques are possible.
- Summary: Manual text selection was slower and more effortful, but participants preferred it and associated it with greater psychological ownership and better perceived performance.The benefit may partly reflect greater control over selection placement, whereas automatic selections tended to be placed at the beginning of the do…
- Summary: Experiment 3 tested six techniques spanning fully AI-written summaries, human–AI exercises, feedback, and fully human-written comments.The techniques included receiving a summary, completing a fill-in-the-blank exercise, writing a prompt, answering a practice question, receiving feedback, and writing a comment without AI.
- Summary: The resulting AI margin note was always a summary, but margin notes can involve humans and AI in other ways.
7 AI Margin Note Techniques
Experiment 3 implemented six AI margin note techniques spanning entirely AI-generated to entirely user-written text. All techniques required manual text selection and pressing a blue “Comment” button, while varying how humans and AI contributed to note creation.
- Technique continuum: The six techniques formed a continuum from fully AI-written summaries to fully user-written comments, with intermediate designs combining generation, prompting, answering, or feedback.All techniques used manually selected text and the same blue “Comment” button.
- Technique continuum: The least-human technique presented an LLM-generated summary of manually selected text, matching the manual text-selection technique from Experiment 2.The summary was entirely written by AI.
- Hybrid techniques: Intermediate techniques required users to complete AI-generated summaries, write custom prompts, or answer AI-generated reading-comprehension questions with feedback.Fill-in-the-blank summaries used one or two dropdown blanks, while practice questions required users to submit answers and correct mistakes.
- Human-authored techniques: The most-human techniques required users to write comment text either with LLM feedback and improvement suggestions or without any AI assistance.The no-assistance condition served as a baseline resembling existing commenting systems.
8 Experiment 3: Human and AI Involvement
Experiment 3 compared six techniques for generating AI margin notes with different levels of human and AI involvement. Participants generally preferred more AI involvement, despite lower psychological ownership, while comprehension differences were only marginal.
- 8 Experiment 3: Human and AI Involvement: Reading comprehension did not significantly differ across techniques: summary scored mdn =4, iqr =2.25 versus mdn =5, iqr =2 for all others.The difference was marginal (𝑝= .054) and not statistically significant.
- 8 Experiment 3: Human and AI Involvement: Participants generally preferred more AI-involved techniques, despite lower psychological ownership and only marginal reading-comprehension differences.Blank ranked first overall, followed by qestion, summary, prompt, feedback, and none.
- 8 Experiment 3: Human and AI Involvement: Faster techniques were summary and blank, which were significantly quicker than all other techniques.Summary had mdn =3.96, iqr =2.78, while blank had mdn =4.92, iqr =2.80.
- 8 Experiment 3: Human and AI Involvement: More human-involved techniques produced greater psychological ownership, with none and feedback highest and summary lowest.None had mdn =100, iqr =4.38; feedback had mdn =97, iqr =9.34; summary had mdn=23.75, iqr=60.88.
- 8 Experiment 3: Human and AI Involvement: More human-involved techniques were more mentally demanding, associated with poorer perceived performance, and more effortful.For Mental Demand, none and feedback scored higher than prompt, blank, and summary.
9 Discussion
The discussion recommends integrated AI margin notes with manual text selection and balanced human–AI involvement, while identifying preferences, ownership, comprehension, transfer, and long-term-use implications. It also notes experimental and interface limitations that motivate future research.
- Integration: Participants preferred integrated AI margin notes because they support convenient text-specific prompting, avoid attention shifts, and provide quick access to prior responses.The authors therefore recommend integration, deictic prompting, reduced interface switching, and persistent retrieval of prompts and responses.
- Selection Automation: Participants preferred manual text selection despite greater time and effort because it encouraged reading, increased psychological ownership, and produced shorter, more uniformly distributed selections.The authors recommend manual selection while adding control over automatically placed notes when automation is necessary.
- Human and AI Involvement: Participants preferred more manual control over selections but more AI involvement in comments, and the most preferred technique combined generated summaries with fill-in-the-blank exercises.That technique was perceived as moderately effortful, moderately ownership-enhancing, and supportive of participants’ perceived performance.
- Human and AI Involvement: More cognitively demanding techniques did not improve reading comprehension, despite prior work linking human involvement and cognitive engagement with comprehension.Additional tests using 24-hour delays, excluding participants with higher background knowledge, and addressing a possible 0-6 ceiling effect found no significant differences.
- Design Implications: The authors recommend balancing human and AI involvement because fully automated summaries were less preferred than techniques retaining some human responsibility, although lower involvement may not harm comprehension.Future designs could explore digital ink and sketches for prompting LLMs.
- Transfer and Limitations: AI margin notes could transfer to code editors and general software help-seeking by persisting context-specific prompts and responses near selected code or relevant interface elements.The approach could also support automatically overlaid notes based on past user behavior, while long-term effects may include hallucination-related overreliance, reduced note-taking ability, and educational harms.
- Transfer and Limitations: Ecological validity was limited by fixed numbers of notes and responses and one-technique-at-a-time interaction, while margin-note space restricts content and interaction rounds.Future work should examine more flexible real-world use and ways to extend notes without making them chat-like.
10 Conclusion
The paper proposes AI margin notes, which integrate LLM capabilities into document text through document-reader commenting features. Across three experiments, participants valued integrated AI margin notes and creating them manually.
- The paper proposes AI margin notes that integrate LLM capabilities into document text through document-reader commenting features.
- Three experiments evaluated AI margin notes across integration, selection automation, and human–AI involvement.
- Participants valued integrated AI margin notes and creating them manually.
A Appendix
The appendix reports statistical test results for Experiments 1–3 and lists pairwise p-values across duration, psychological ownership, mental demand, performance, effort, and text similarity.
- Statistical test results for Experiments 1, 2, and 3 are provided in Tables A.1, A.2, and A.3, respectively.
- Pairwise comparisons are reported across duration, psychological ownership, mental demand, performance, effort, and text similarity.
- For example, the none–blank comparison has p-values of .003, < .001, < .001, .005, < .001, and < .001 across the six listed measures.