Source-linked AI summary
Provenance Before Prose: Claim-Locked Reporting
Xiao Fan, Jingyuan Li, Hongbin Guo, Yubo Han, Yi Zhang
TL;DR
LLM statistical reports can distort evidence-bearing content even when source results are fixed, making reproducibility and claim control important. The paper fixes each reportable claim’s provenance, numbers, direction, and language strength before generation, then lets the LLM write only connective prose. Across fMRI and RCT reporting, claim-locked reporting improves reproducibility over a hybrid template by 37.4 and 20.5 points, respectively.
Problem
LLM statistical reports may drift numbers, reverse directions, or overstate claims despite using source-supported content, so evidence-bearing meaning must remain stable across runs.
Method
Claim-locked reporting binds each reportable claim to its evidence source, numbers, direction, and allowed language strength before deterministic rendering and connective prose generation.
Results
37.4 and 20.5 points are the reproducibility improvements over the hybrid template in fMRI and RCT reporting, respectively.
Takeaways & Limitations
Claim-locked reporting fixes which evidence-bound claims are renderable before generation, improving cross-run stability while governing rhetorical strength.
Takeaways & Limitations
Claim-locked reporting does not perform statistical validation or guarantee truth beyond the upstream evidence record, and reproducibility alone can still yield incorrect reports.
Abstract
from arXiv · showhide
Large language models (LLMs) can fluently verbalize statistical evidence, yet statistical reports can still drift numerical values, invert effect directions, or restate thresholded contrasts as categorical effects. We frame these failures as a control problem: the evidence-bearing content of a scientific report should be fixed by structured statistical results rather than sampled during prose generation. We therefore use cross-run reproducibility to stress-test whether report-visible numbers and claims are bound before prose generation. Existing controls operate at the text or slot level; a deterministic hybrid template reproduces only 61.1% of report-visible numerical content across seeds because the LLM still selects which findings and numbers the template renders. We propose claim-locked reporting, a provenance-before-prose protocol that fixes the evidence source, numbers, direction, and allowed language strength of each reportable claim before the LLM writes only connective prose. Across fMRI functional-connectivity reporting and randomized controlled trial reporting on Evidence Inference 2.0, claim-locked reporting improves reproducibility over the hybrid template by 37.4 and 20.5 points, respectively. Blinded human audits support the observed direction-preservation and governance trends. In an fMRI cost analysis with DeepSeek, claim-locked reporting also yields the lowest observed token use and median generation latency.
1. Introduction
LLM-generated statistical reports can change evidence-bearing content across runs even when the underlying results are fixed. Claim-locked reporting addresses this gap by fixing reportable claims before prose generation and improves reproducibility across fMRI and RCT reporting.
- Statistical claim distortion includes numerical drift, direction reversal, and overstated inferential strength despite source-supported wording.
- Existing text-level controls constrain generation or output form, but leave the LLM to choose which statistical inferences to state and how strongly.
- Cross-run reproducibility is used as a stress test of whether evidence-bearing report content is fixed before prose generation.
- Claim-locked reporting binds each reportable claim to its evidence source, numbers, direction, and allowed language strength before the LLM writes connective prose.
- 37.4 and 20.5 points are the reported reproducibility improvements over the hybrid template in fMRI and RCT reporting, respectively.
2. Related Work
Prior data-to-text and constrained-generation approaches separate or constrain content access, structure, and realization. Statistical reporting additionally requires preserving numerical values, directions, inferential scope, and interpretive strength.
- Classical natural language generation separates content planning from surface realization, whereas neural data-to-text systems often learn content selection and realization jointly.
- Retrieval, constrained decoding, structured generation, and post-hoc verification constrain evidence access, output structure, or generated text.
- Statistical reporting requires preserving numerical value, effect direction, inferential scope, and interpretive strength in addition to source grounding.
3. Claim-Locked Reporting
Claim-locked reporting converts fixed statistical evidence into ledger-bound claims, audits their risks, and renders locked content before the LLM writes connective prose. Its evaluation separately tests stability, governance, and correctness-related properties.
- 3.1 Statistical Evidence Record: The evidence record contains already-computed statistical results and serves as the fixed input to claim construction rather than raw imaging or trial documents.
- 3.2 Locked Statistical Claim: The claim ledger stores typed claims with evidence pointers, numerical fields, effect direction, support level, and allowed language strength.
- 3.2 Locked Statistical Claim: The risk auditor checks numerical support, direction preservation, entity/source support, inferential scope, and interpretive strength using structured evidence fields.
- 3.2 Locked Statistical Claim: The policy controller can only downgrade flagged claims’ language strength and excludes forbidden claims from report blocks.
- 3.3 Report Generation: The deterministic renderer emits numerical content from evidence pointers, while the LLM writes only connective paragraphs under policy constraints.
- 3.5 Evaluation Axes: The evaluation reports cross-run reproducibility, claim governance, and per-run numerical and directional correctness as separate axes.
4. Experimental Setup
The study evaluates claim-locked reporting on cohort-level neuroimaging and clinical-trial reporting under matched evidence records, providers, and seeds. Baselines vary structural or post-hoc controls, while the hybrid isolates slot-level control.
- The experiments cover cohort-level neuroimaging reports and clinical-trial summaries that stress different parts of the same statistical-reporting control problem.
- fMRI FC setting: The fMRI setting uses obesity-related functional-connectivity reporting, where covariates and confounds require preserving claim scope after adjustment.
- RCT setting: The RCT setting samples 200 consensus-labeled Evidence Inference 2.0 records with non-null directions and at least two numerical tokens per evidence span.
- Methods and fairness: Five grounded baselines combine evidence injection with free-form, prompt, structured, retrieval, or post-hoc controls; the hybrid adds deterministic rendering while retaining LLM slot selection.
- Main fMRI comparison: Table 1 compares seven fMRI methods across 20 cells per row using reproducibility, unhedged strong statements, and numerical audit flags.
- Methods and fairness: All methods receive identical structured evidence records for each cohort, provider, and seed, differing only in the control point between record and report.
5. Results
Claim-locked reporting makes evidence-bearing content far more stable across runs than controls that leave content selection to the LLM. Results also show separate gains in governance, inferential-scope protection, numerical validity, human-audited direction preservation, and operational efficiency.
- Cross-run reproducibility: 98.5% cross-seed reproducibility was achieved by claim-locked reporting, compared with 61.1% for the hybrid template and 15.2–32.3% for grounded baselines.The claim-locked advantage over the hybrid template was 37.4 points, with a 95% CI of [+15.1, +59.7].
- Governance and numerical flags: 2.25 unhedged-strong sentences for the hybrid template fell to 0.75 for claim-locked reporting, while both methods had 0.00 per-run numerical flags.The paired difference for unhedged-strong sentences was −1.50, with a 95% CI of [−2.05, −1.00].
- Operational efficiency: Claim-locked reporting used one LLM call and had the lowest observed input/output token use and median latency in the DeepSeek fMRI comparison.The post-hoc verifier used two LLM calls; deterministic components added no LLM calls and had small local runtime relative to generation latency.
- Group-covariate absorption: 10,041 FDR-surviving group-effect edges fell to 8 after BMI adjustment in the in-house cohort, while the HCP cohort fell from 550 to 0.Claim-locked binds the before-and-after-adjustment claim so the primary count cannot be rendered with independent-effect wording.
- RCT transfer and numerical validity: Claim-locked reporting produced no harmful numerical-result fabrications, compared with approximately 2.5% for the hybrid template and 5.9% for free-form generation.The calibrated audit separates parsing artifacts and benign metadata from harmful statistical-result fabrications.
- Human audits: Claim-locked reporting produced no direction inversions, whereas baseline inversion rates ranged from 5.6% to 14.3% in the RCT direction audit.The audit used a stratified subset with 98.0% raw agreement and Cohen’s κ = 0.970.
6. Conclusion
Claim-locked reporting fixes reportable statistical claims before prose generation, leaving the LLM to write only connective prose. It improves reproducibility and supports governance while reducing writer-side token use and median latency in the fMRI analysis.
- 61.1%-to-98.5% reproducibility: fixing evidence-bound claims adds stability beyond deterministic rendering of LLM-selected slots.
- Claim-locked reporting fixes reportable claims, sources, directions, numbers, and language strength before writing, while the LLM supplies only connective prose.
- In the RCT setting, manual calibration found no harmful statistical-result fabrication in claim-locked outputs, whereas slot-level rendering still admitted unsupported results.
- Claim-locked reporting also reduced writer-side token use and median generation latency in the fMRI cost analysis.
Limitations
Claim-locked reporting governs how structured evidence is verbalized but does not validate upstream statistics or guarantee truth beyond the evidence record. Its evaluation is bounded by predefined, benchmark-specific risks, targeted audits, and limited RCT direction labels.
- Claim-locked reporting does not perform statistical validation, discover mechanisms, or guarantee truth beyond the evidence record.
- Upstream preprocessing, model specification, covariate selection, or evidence-extraction errors propagate into the report if already present.
- Cross-run reproducibility indicates stable evidence-bearing content when it passes, but not correctness; separate correctness and governance measures remain necessary.
- Automatic metrics test predefined risks rather than replacing expert judgment, and the risk taxonomy is benchmark-specific rather than exhaustive.
- The RCT evaluation covers only increased and decreased direction labels; null or inconclusive findings require a neutral ledger state forbidding efficacy-implying language.
- The institutional FC cohort cannot be redistributed under its data-use protocol, so public HCP and Evidence Inference 2.0 settings are included.
Ethics Statement
The work uses institutional and public datasets under stated governance conditions and releases implementation materials where permitted. Its claim ledger records provenance, numerical fields, direction, risk tags, and language constraints for report generation.
- The institutional FC cohort was collected under IRB approval with written informed consent and cannot be redistributed under its data-use agreement.
- The public HCP S1200 cohort is used under HCP Open Access Data Use Terms, while Evidence Inference 2.0 is a public clinical-trial-literature benchmark.
- The claim ledger is the single object passed from pregeneration control logic to the writer, with evidence pointers, numerical fields, direction, support level, risk tags, and language limits.
- The worked BMI-absorption ledger entry records source edgewise_glm_summary.M1_to_M3, 10,041 primary edges, 8 BMI-adjusted edges, 99.9% collapse, and cautious language.
- The renderer emits ledger-bound numbers and direction, while forbidden-language constraints are passed to the writer as hard restrictions.
A.2 Automatic evaluation metrics
The evaluation uses separate axes for reproducibility, rhetorical governance, and per-run numerical or directional correctness. Automatic checks compare report content with structured evidence, detect strong language and risk-conditioned hedging, and distinguish control configurations by which reporting decisions remain under LLM control.
- Automatic metrics: Num matches report-visible numbers against evidence JSON leaves, an explicit derived-metric whitelist, and a small conventional-constant whitelist.
- Automatic metrics: Num uses absolute tolerance 10−3 or relative tolerance 5%, filters common section and ordinal numbers, and deliberately excludes arbitrary pairwise ratios.
- Automatic metrics: Strong counts sentences containing strong assertion terms without a hedge, using fixed lexicons established before evaluation.
- Automatic metrics: Hedge measures cautionary language only for risk-flagged claims, testing whether predefined caution constraints reach the report surface.
- Control configurations: The FC lexical direction check is supplementary because negation and comparator phrasing make it brittle, while methods differ in which reporting decisions remain under LLM control.
- Control configurations: Claim-locked reporting binds source, numbers, direction, and permitted language strength before prose generation, unlike hybrid rendering of LLM-selected slots.
C.1 Component analysis
The component analysis separates deterministic realization from claim auditing and policy control. Deterministic rendering produces most of the reproducibility gain, while the full controls additionally reduce unhedged strong language.
- Builder only leaves numerical realization to the LLM, whereas Builder + renderer isolates deterministic realization by disabling the risk auditor and policy controller.The configurations are evaluated under the same records, providers, and seeds.
- Reproducibility rises from 85.0% to 98.0% for fMRI and from 90.3% to 100.0% for RCT when deterministic rendering is added.
- Full claim-locked reporting leaves fMRI reproducibility essentially unchanged at 98.5% while reducing Strong from 1.40 to 0.75.The full configuration activates the risk auditor and policy controller in addition to deterministic rendering.
- The shared auditor interface uses five control targets across domains: numerical support, direction preservation, entity or source support, inferential scope, and interpretive strength.The concrete evidence tags differ between fMRI and RCT settings.
E. Human Audits
Blinded human audits provide reader-facing checks on the automatic governance and direction-preservation measures. Reviewers found claim-locked outputs lowest in violation counts, and independent RCT labeling showed near-perfect agreement.
- The blinded fMRI audit used four independent reviewers who annotated matched snippets without seeing method identity.They counted unhedged strong language and unsupported content violations.
- Claim-locked outputs have the lowest average counts for both unhedged strong language and unsupported content in the blinded fMRI audit.Hybrid is intermediate for strong language but highest for unsupported content, while free-form is highest for strong language.
- RCT direction auditing classified reports as preserved, inverted, or unresolved relative to the gold direction, excluding unresolved reports from the inversion-rate denominator.The audit covered all 840 generated reports in a 30-record subset.
- A second reviewer achieved 98.0% raw agreement and Cohen’s κ = 0.970 on a stratified 50-report subset.The single disagreement was between preserved and inverted.
F. RCT Numerical-Flag Calibration
The RCT numerical-flag calibration distinguishes benign or regex-related flags from harmful statistical-result fabrication. This changes how raw numerical-flag counts should be interpreted across reporting methods.
- Raw Num is a high-recall numerical flag that does not distinguish unsupported statistical results from benign mismatches.
- Manual calibration classifies each numerical flag as Regex FP, a benign metadata or support-window mismatch, or harmful statistical-result fabrication.Harmful fabrication means a report-visible statistical value is absent from the evidence record.
- Claim-locked has the largest raw flag count, but all of its flags are benign or false positives, with no harmful statistical-result fabrication.
- The hybrid template produces unsupported statistical results in 2.5% of outputs despite having fewer raw flags.This occurs because the LLM still selects the slot content.