Source-linked AI summary
Balancing Evidence and Interpretation: Historical Grounding Ratio as a Design Parameter for AI-Generated Urban Storytelling
Fuyang Zhang, Maurice Benayoun
TL;DR
Location-aware generative systems lack an operational way to measure how multiple relevant sources compose the final narrative. This paper introduces HGR, implements controlled evidence allocation in GeoDrama, and evaluates it in an 18-participant within-subject walking study. Higher HGR increased perceived narrative–place relevance, whereas four other experiential measures were highest in the balanced condition.
Problem
Existing systems do not clearly represent or measure how much content from each jointly relevant source appears in the generated output.
Method
The paper defines HGR as the proportion of claim-bearing delivered information units supported by historical archives and tests three controlled allocation conditions in GeoDrama.
Results
Higher HGR strengthened perceived narrative–place relevance, while historical understanding, scene integration, information appropriateness, and exploration intention were highest in the balanced condition.
Takeaways & Limitations
HGR is a design parameter for comparing and adjusting source-allocation strategies according to particular experiential goals, not a metric to maximize universally.
Takeaways & Limitations
The study cannot establish a universally optimal historical-grounding level.
Abstract
from arXiv · showhide
Location-aware generative systems can now select historical archives and real-time contextual information based on a user's surroundings to automatically generate narratives for urban heritage walks. Yet when multiple sources jointly inform generation, existing systems provide neither a clear representation of how much content from each source actually appears in the output nor an operational means of measuring it. We introduce the Historical Grounding Ratio (HGR), defined as the proportion of claim-bearing information units in a generated narrative that are supported by historical archives. HGR turns the realized share of historical evidence in a narrative into a directly measurable design parameter. In GeoDrama, a mobile narrative system, we created three conditions that used a common retrieval procedure and comparable evidence-bundle sizes while varying the allocation of information from different sources during generation. We evaluated how changes in HGR affected narrative experience through a within-subject walking study with 18 participants. Increasing HGR significantly strengthened the perceived relevance between narrative content and the specific location. However, historical understanding, integration with the visible scene, appropriateness of the amount of information, and intention to explore further did not increase monotonically with HGR; all four measures were highest in the intermediate, balanced condition. These findings show that designing location-aware generative interfaces involves not only retrieving relevant material but also determining how information from different sources composes the final output. HGR offers an operational measure for comparing information-allocation strategies and their experiential consequences.
1 Introduction
Location-aware generative storytelling must decide not only which materials are relevant and supported, but how historical and situated information compose the delivered narrative. HGR operationalizes this composition, and GeoDrama tests its experiential effects across three evidence-allocation configurations.
- Design problem: Evidence allocation determines how a limited output budget is divided between historical evidence and content connected to the surrounding environment.More historical content may provide concrete facts, while more situated content may strengthen immediate connection but reduce historical information encountered.
- HGR: Historical Grounding Ratio (HGR) measures the realized proportion of archive-supported information units among all archive-supported, scene-anchored, and interpretive or bridging units.It measures the delivered narrative rather than the ratio specified in a prompt and is distinct from retrieval relevance and claim correctness.
- Implication: HGR is best treated as an adjustable interface-policy parameter for particular experiential goals rather than a quality metric to maximize universally.The paper presents HGR as a common scale for comparing allocation strategies and their trade-offs.
- GeoDrama: GeoDrama uses location-constrained archival retrieval, live-scene sensing, candidate reranking, and controlled narrative generation to manipulate evidence allocation.Its three conditions share retrieval, model, generation, structure, candidate-evidence amounts, and output capacities, while varying historical and situated information quotas.
- Study: 18 participants completed a counterbalanced within-subject walking study in which each experienced all three conditions, yielding 54 walking segments.The study measured realized HGR and five experiential outcomes during a fixed-route walk through a historically dense urban area.
- Findings: Increasing HGR significantly strengthened perceived narrative–place relevance, while the other four experiential measures were highest descriptively in the balanced condition.Historical understanding, visible-scene integration, information appropriateness, and intention to explore further did not improve in parallel with HGR.
2 Related Work
Research on location-aware and generative interfaces has advanced contextual sensing, content selection, and factual support, but has paid less attention to how multiple relevant sources compose the final narrative. This section positions source composition after retrieval as a distinct content property requiring separate measurement.
- Location-aware mobile narratives: Earlier location-aware systems selected pre-authored content based on location, while later work adapted triggering, timing, pacing, and route-based presentation.These systems generally determined the information within each narrative segment during content production.
- Context awareness: Context-aware interfaces expanded content selection beyond location to include user interests, environmental conditions, route characteristics, cameras, and augmented reality.These signals can connect immediate surroundings with supplementary information and support continuous selection during a walk.
- Generative interfaces: Generative interfaces allow historical materials, live scenes, and other contextual information to participate directly in runtime content generation.This changes the final narrative from a predetermined content property into an outcome formed during generation.
- Situated intelligent interfaces: Recent systems jointly use multimodal context, user state, task progress, interaction history, and external knowledge to generate responses suited to current situations.Examples include wearable, smartwatch, augmented-reality, social-interaction, and camera-based assistance systems.
- Source composition: The paper focuses on measuring the realized composition of historical, situated, and other information in generated narratives rather than redefining retrieval quality, factual consistency, or source attribution.It treats post-generation source composition as a content property that must be measured separately.
- Information use: Retrieved information does not enter generated output uniformly because models may select, compress, reorganize, or transform materials during generation.Relevant context placement and organization affect use, and source composition cannot be inferred directly from retrieval counts or prompt targets.
3 Research Questions
The paper asks how historical evidence share can become a measurable and controllable design variable distinct from retrieval relevance and attribution, and how different realized levels affect walking experiences. It motivates these questions through limited attention, model transformation of inputs, and the underexamined proportions in which relevant sources enter delivered content.
- Gap: Existing work emphasizes contextual input, retrieval relevance, and factual attribution, while giving less attention to the proportions in which multiple relevant sources enter delivered content.The paper frames this as an output-side evidence-allocation problem between retrieval and user experience.
- Operationalization: The paper introduces HGR to measure the realized share of historical archives in situated generative narratives.HGR is defined over claim-bearing information units in the delivered text that are supported by historical material.
- Motivation: Mobile users have limited interaction time and attention, so systems cannot present all retrieved materials in full.Generative models may select, compress, and transform inputs, making the historical proportion in delivered content impossible to infer from retrieval counts or prompt targets.
- Research questions: RQ1 asks how the share of historical evidence in a delivered narrative can be operationalized as a measurable, verifiable, and controllable design variable.The variable should remain distinct from retrieval relevance and generation attribution.
- Research questions: RQ2 asks how different levels of historical grounding affect users’ experiences during real-world urban walks when candidate evidence and other generation conditions are held constant.This isolates evidence allocation as the principal experimental difference.
4 Historical Grounding Ratio
The Historical Grounding Ratio (HGR) measures the realized share of archive-supported information units in a delivered narrative, rather than retrieved records or prompt targets. It operationalizes source composition while distinguishing archival, scene-grounded, and interpretive content within finite narrative capacity.
- Definition and Calculation: Narratives are segmented into independently assessable information units, while non-claim connectors and transitions are excluded from the denominator.Category C content includes address terms, grammatical connectors, and transitions that do not form independent claims.
- Source Categories: Archive-supported, scene-grounded, and interpretive units respectively capture historical evidence, current environmental support, and connective or associative meaning.An archival claim remains category A when the current scene also corroborates it, if archival evidence is required to establish the claim.
- Definition and Calculation: HGR is the proportion of archive-supported information units among all valid narrative information units.It is measured from the delivered narrative after generation.
- Interpretation: HGR describes narrative source composition rather than overall quality, so a higher value is not inherently better.Higher HGR leaves more finite narrative space for historically supported information, while lower HGR leaves more space for scene or interpretive organization.
- Experimental Conditions: The study manipulated low, medium, and high historical-grounding conditions while holding retrieval mechanisms, comparable evidence amounts, and narrative capacity broadly consistent.The primary difference was how archival and situated information were selected and organized into the final narrative after retrieval.
- Measurement Consistency: Machine annotations closely matched human HGR judgments, with Pearson r= .984 and ICC(A, 1) = .965, while showing a small negative bias of M= −.037 and MAE = .053.The reported bias weakened rather than strengthened the experimental manipulation.
5 GeoDrama System
GeoDrama is a mobile research prototype that senses location and streetscape context, retrieves geographically eligible historical records, and generates and delivers auditable situated narratives. Its controlled pipeline separates retrieval from post-retrieval information allocation while preserving comparable conditions for experimentation.
- System Architecture: GeoDrama combines location sensing, scene understanding, historical retrieval, target allocation, narrative generation, mobile delivery, and audit logging in one pipeline.The system streams generated sentences as speech and records research events and feedback for auditing.
- Location and Scene: Location determines geographically eligible historical records, while scene descriptions rerank those records and provide context for narrative generation.The visible scene therefore participates in both evidence selection and narrative organization.
- Information Allocation: All experimental conditions use the same sensing and retrieval mechanisms but vary how archival, situated, and interpretive content enters the final narrative.Information allocation is manipulated after retrieval rather than by changing the candidate-generation process.
- Spatial Retrieval: The system uses spatial containment rather than distance radii because historical places have irregular, overlapping extents that need not match present-day addresses.Every polygon containing the current location can activate associated historical records for retrieval.
- Evidence Retrieval: The evidence bundle contains ten reranked historical records totaling 776–886 Chinese characters, with M≈830.Candidates are first geofenced and then semantically reranked within that constrained pool.
- Retrieval Performance: Scene-based semantic reranking succeeded in 43 of 54 segments with a mean latency of 9.44 ms, changing the ten-record bundle order in every successful segment.For the remaining 11 segments, retrieval used deterministic random ordering within the geofenced candidate pool because no runtime scene description was available.
- Narrative Generation: The generator receives location-specific evidence, a scene description, a fixed nine-sentence scaffold, and an allocation instruction under identical model and decoding settings.Across 54 final segments, mean generation latency was 4519 ms, ranging from 3309–11366 ms.
6 Field Study
The field study tested three historical-grounding configurations in GeoDrama during a counterbalanced walking study in Nanjing. Eighteen participants received situated narratives while completing all three conditions.
- Participants and setting: 18 participants completed a within-subject walking study in a historical district of central Nanjing.All participants completed the three conditions and were included in analysis.
- Experimental design: Each participant experienced situation-dominant, balanced, and evidence-dominant conditions during one continuous walk.The three route segments used complete counterbalancing across condition order and segment assignment.
- Measures: Post-segment ratings, retrospective comparisons, and semi-structured interviews assessed place relevance, historical understanding, scene integration, information fit, and further-exploration intention.The measures were collected after each route segment and supplemented by qualitative interpretation of the quantitative patterns.
- Participants and setting: The study area contained 161 points of interest and 161 overlapping, nested historical polygons aligned with the experimental route.A single observed location could correspond to multiple historical places at different spatial scales.
- Experimental design: The conditions varied the allocation of archive-supported and in-situ information while using the same retrieval mechanism.The balanced condition used a relatively even configuration, whereas the evidence-dominant condition allocated more narrative content to archive-supported material.
- Procedure: Participants received location-triggered narratives through a mobile interface while walking independently along three consecutive route segments.Researchers demonstrated operation beforehand and intervened only for location, network, or interface failures.
7 Results
The manipulation produced distinct historical-grounding levels without materially changing narrative quantity. Higher grounding strengthened place relevance, while the remaining experiential outcomes favored the balanced condition.
- Manipulation check: .293, .601, and .754 were the mean segment-level HGRs for situation-dominant, balanced, and evidence-dominant conditions, respectively.The condition effect was significant, χ2(2) = 31.00, p < .001, and all corrected pairwise comparisons were significant.
- Source composition: As HGR increased, archive-only and archive-and-scene-supported units increased while scene-only and interpretive units decreased.The manipulation therefore changed source composition within a comparable amount of discourse rather than simply increasing content volume.
- Source composition: 264, 254, and 267 valid information units occurred across the situation-dominant, balanced, and evidence-dominant conditions, respectively.The conditions averaged 14.67, 14.11, and 14.83 units per segment, and each narrative contained nine sentences.
- Experience ratings: Historical understanding, scene integration, information fit, and intention to explore further did not show significant corrected pairwise differences.All four measures nevertheless had their highest means in the balanced condition.
- Composite outcomes: The balanced condition’s Q2–Q5 composite was 5.40, compared with 4.72 and 4.97 in the situation-dominant and evidence-dominant conditions.An exploratory quadratic contrast found the balanced condition exceeded the mean of the two extremes by 0.56 scale points.
- Experience ratings: Place relevance increased successively with HGR, with the evidence-dominant condition significantly exceeding the situation-dominant condition.The evidence-dominant versus situation-dominant effect was stronger for Q1 than for the average of Q2–Q5.
- Retrospective comparisons: Retrospective choices separated overall preference and information fit from connection to place.The balanced condition received the most selections for preference and information amount, whereas the evidence-dominant condition received the most selections for connection to place.
8 Discussion
The discussion interprets HGR as a measurable interface policy rather than a universal quality dial. Higher grounding favored content–place relevance, while a balanced configuration better supported several broader experiential goals.
- Mechanisms: The study’s source-composition manipulation mainly replaced scene-only and interpretive units with archive-supported units while keeping unit and sentence counts similar.This supports interpreting the experience differences as related to supporting sources, not merely to narrative quantity.
- Interpreting the findings: Increasing HGR strengthened perceived association between narrative content and the specific place but did not improve all experiential dimensions in parallel.The authors characterize the pattern as a dissociation among experiential outcomes rather than a uniform response curve.
- Design implications: Higher grounding appears advantageous when the design goal emphasizes relationships between historical facts and a specific place.For integrated experiences involving understanding, scene integration, information fit, and continued exploration, the discussion motivates balanced configurations.
- Mechanisms: Historical evidence may support scene understanding by providing identifiable objects and temporal references that organize interpretation of the visible environment.The discussion proposes that historical grounding need not compete with scene integration when it supplies locatable, time-situated references.
- Contribution: HGR turns historical source contribution from an implicit generation outcome into a measurable, manipulable, and auditable interface-level design variable.The contribution distinguishes source allocation from retrieval relevance and generation attribution.
- Contribution: Interface policies should identify the experience to optimize rather than assume that more historical evidence is always better.The authors frame HGR as a design policy configured according to goals, not as a quality score with a universally optimal maximum.
- Boundaries and future work: The apparent advantage of an intermediate HGR requires further testing across experiential goals and settings.The discussion explicitly proposes examining whether this advantage is stable and whether different goals require different configurations.
9 Conclusion
The conclusion presents HGR as a measurable, manipulable parameter for controlling how historical, situated, and interpretive content composes generated narratives. The study finds that higher HGR strengthens narrative–place relevance, while a balanced configuration better supports several other experiential outcomes.
- Contribution: HGR measures the proportion of claim-bearing information units in delivered text supported by historical archives, rather than the amount of historical material retrieved.It makes source contribution a measurable design object focused on the narrative users actually experience.
- Contribution: The within-subject field study showed that different source configurations produced clearly separated achieved grounding levels in generated texts.This indicates that HGR can be manipulated through source allocation strategies, although achieved values must be measured in the delivered text.
- Findings: Increasing HGR strengthened participants’ perceived association between narrative content and a specific place.The effect concerns narrative–place relevance rather than a uniformly improved walking experience.
- Findings: Historical understanding, scene integration, information fit, and intention to explore further did not improve in parallel with HGR; the balanced configuration offered an advantage across their combined pattern.The study therefore reports a dissociation among experiential outcomes rather than a universal optimum.
- Implications: HGR should be configured and validated for a specific experiential goal, not treated as a monotonic generation-quality score.Higher grounding suits emphasizing place relationships, whereas integrated experiences may require balance among historical, situated, and interpretive content.
- Implications: Prompt quotas cannot be assumed to equal achieved grounding because generative models may deviate from source requirements in their final expression.Systems depending on a particular HGR should measure the realized value in delivered text rather than infer it solely from prompts or generation settings.
Ethics Statement
The study received institutional ethics approval, and all adult participants volunteered and provided electronic informed consent. Participants were informed about possible narrative errors and their right to withdraw.
- Institutional ethics approval was obtained before data collection.
- All participants were adults, participated voluntarily, and provided electronic informed consent.
- Participants were warned that generated narratives might contain errors and could withdraw from the study.
GenAI Usage Disclosure
Generative AI supported both the GeoDrama system and research tasks, while the authors retained responsibility for verifying the study and manuscript.
- Generative AI was studied in GeoDrama, whose evaluated system used Qwen/Qwen3-8B for narrative generation.
- Claude and OpenAI Codex assisted with language revision, analysis scripting, debugging, and figure production.
- The authors reviewed and verified study decisions, data collection, statistical results, interpretations, citations, and manuscript content.
- The authors accepted full responsibility for the accuracy and integrity of the submitted work.