Source-linked AI summary

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Miao Liu, Zhizhe Liu

arXiv:2608.24842v1cs.CLcs.AI

TL;DR

The paper asks whether information retrieved by AI analysts actually influences their investment judgments, rather than merely being available to report. It separates retrieval from decision influence in controlled long-context experiments and finds that integration can fall to the empirical noise floor despite accurate retrieval. The authors therefore distinguish reading from using and bound their mechanism claims to AI processors.

  • Problem

    Finding information and using it are distinct stages of decision support, but whether retrieved information influences investment judgment remains the central gap.

  • Method

    The study separately measures retrieval and decision influence while varying unrelated context and examining memory and workflow mechanisms.

  • Results

    A risk disclosure’s decision influence reaches the empirical noise floor through 128,000 tokens even though the model retrieves it correctly, with the pattern extending beyond one model family.

  • Takeaways & Limitations

    Reading and downstream decision influence can separate sharply, so an AI analyst’s effective decision context need not equal its advertised context window.

  • Takeaways & Limitations

    The mechanism claims are bounded to AI processors rather than human decision-makers, and some model families show false positives on neutral filings.

Abstract

from arXiv · show

Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.

1 Introduction

The paper identifies a retrieval–integration gap: long-context models can retrieve decision-relevant disclosures accurately while failing to use them in investment judgments. Controlled experiments show that model capability delays this failure, but workflow architecture determines whether information reaches the decision.

  • Core finding: 3.2 percentage points: the primary model’s sell probability rises by this amount with a risk disclosure in a 2,000-token filing.Between 8,000 and 32,000 tokens, the disclosure’s influence becomes indistinguishable from neutral-text effects and remains at that noise floor through 128,000 tokens.
  • Retrieval versus use: At 128,000 tokens, the primary model retrieves the disclosure for all twelve firms without false retrievals on neutral filings.Retrieval remains accurate even as downstream decision influence disappears.
  • Robustness: Across three independently trained model families, the retrieval–integration gap recurs, although its binding context length differs across families.In complete filings, the real disclosure has essentially no influence despite correct retrieval in all twenty cases.
  • Capability: The largest open-weight model retains a 3.4-percentage-point effect at 128,000 tokens, approximately matching the mid-sized model’s 2,000-token effect.Capability shifts the failure frontier outward but does not establish immunity; the effective decision context need not equal the advertised context window.
  • Mechanism: Representation analyses show that disclosure content remains encoded where read, while decision-position probes lose disclosure-specific legibility as context grows.Memory interventions indicate that compressed summaries and attention-based source lookup jointly transmit disclosures into judgments, with both channels weakening under competing context.

2 Background and Hypotheses

The paper reframes AI financial analysis as a delegated, multistage process in which retrieval and judgment are separate transmission stages. It asks whether a model that can retrieve a disclosure actually integrates it into a decision, and motivates evaluation against consequential judgment outcomes rather than retrieval proxies alone.

  • Delegation: Delegated AI artifacts must carry information across reading, internal representation, and judgment stages that are not directly observable to users.A verifiable intermediate operation can therefore be completed without completing the delegated objective.
  • Conceptual framework: The retrieval–integration gap is the divergence between what an AI analyst can retrieve from a document and what affects its judgment.The paper measures this divergence rather than inferring it from a single retrieval score.
  • Processing costs: The framework treats information acquisition and integration as separate costs that need not fall together as AI reduces retrieval costs.This makes the retrieval–integration distinction operational for machine readers.
  • Evaluation gap: Retrieval accuracy is a convenient, observable proxy that does not establish whether retrieved information affects the supported decision.This separates machine readability from machine decision usefulness.
  • Hypotheses: The paper holds focal-firm information fixed while increasing unrelated context and predicts that disclosure influence declines as context grows.The hypotheses distinguish context-driven integration failure from retrieval failure and propose workflow reorganization as a remedy.

3 Experimental Design

The design isolates how much a single known risk disclosure changes an AI analyst’s judgment as surrounding context grows, while measuring retrieval separately. It uses controlled target-versus-neutral filings, multiple decision readouts, and context-placement and ecological extensions to distinguish retrieval from downstream use.

  • Core design: The experiment measures a disclosure’s marginal influence on investment judgment as unrelated filing text increases from 2,000 to 128,000 tokens.The focal firm and target risk remain fixed while matched accounting and legal prose expands the context.
  • Measurement: Retrieval and decision influence are measured in separate deterministic calls, making their divergence an observed retrieval–integration gap.The retrieval call asks what the filing says, while the decision call measures movement in the delegated judgment rather than intermediate retrieval accuracy.
  • Measurement: Marginal decision influence compares the judgment with the disclosure against the average judgment across five neutral replacements.Averaging neutral replacements absorbs idiosyncratic shifts and keeps the measure tied to the delegated task’s output units.
  • Core design: Each firm has one target filing with a firm-specific risk paragraph and five otherwise identical neutral replacements.The counterfactual comparison changes only the paragraph containing the disclosure.
  • Context manipulations: The study tests whether context arrangement matters by placing filler before or after the focal filing and by varying disclosure distance at fixed total lengths.The before/after comparison is the pre-specified secondary test of H3; distance is varied at 32,000 and 128,000 tokens.
  • Extensions and validity: Extensions replace unrelated filler with other firms’ risk disclosures, insert targets into edited complete 10-K filings, and remove authored text in an additional design.The ecological condition accepts a changed firm information set in exchange for realism, while authored candidates undergo coherence auditing.

4 The Retrieval–Integration Gap Under a Fixed Infor-

As unrelated context grows, retrieved risk disclosures increasingly stop affecting investment judgments, despite stable direct retrieval. This retrieval–integration gap replicates across models, tasks, and real filings, while severity ordering can persist even as response magnitude collapses.

  • Replication: Across three model architectures, disclosure use declines from strong short-context effects toward zero as unrelated text accumulates.The focal firm’s information is held fixed, so the decay is not rational updating on new firm information.
  • Severity sensitivity: At 128,000 tokens, response range falls from 0.23 to 0.04 while pairwise severity-ordering accuracy declines only from 0.93 to 0.79.Long-context judgments can preserve ordinal ranking while flattening economic magnitude.
  • Evaluation implication: Retrieval accuracy is therefore an insufficient proxy for decision usefulness because systems can state retrieved facts while their investment judgments remain invariant.The paper identifies this divergence as the retrieval–integration gap.
  • Real filings: Complete real filings show retrieval completeness of 0.931 but approximately zero measured use, reproducing the gap in investors’ information environment.Removing a real disclosure from a roughly 2,000-token excerpt produces a +0.037 average judgment shift, while removal from the complete filing produces +0.000.

5 Mechanism: Where the Disclosure Is Lost

The paper distinguishes reading, transmission, and judgment failures by measuring internal representations at the disclosure and decision positions. Its instruments show that long-context decay occurs after local comprehension, during transmission into the final judgment.

  • Failure accounts: The study separates three possible failures: not forming a usable representation, failing to carry it forward, or failing to weight it.These correspond to reading, transmission, and judgment failures with different design implications.
  • Two instruments: Vocabulary probes test readable words, while sparse-autoencoder features test whether risk concepts are present at selected positions.The instruments measure different internal properties independently.

Appendix C.

Causal interventions implicate two transmission channels—the recurrent running summary and attention-based lookup—in carrying disclosures into judgments. Their capacity-sensitive decay suggests workflow routing, not merely model computation, is central to long-context use.

  • Channel interventions: Attention removal produces 44 to 48 percent of the disclosure’s contrast across ten to eleven of twelve firms.A sham intervention has no directional effect, distinguishing the real attention effect from generic perturbation.
  • Channel interventions: The running summary removes about two-thirds of disclosure influence, while recreating the summary restores nearly half.The removal and recreation effects are approximately 65 percent and 44 percent, respectively.
  • Dual carriers: The running summary and attention lookup each carry substantial influence, with no statistically resolvable ranking between them.Their effects need not sum to one because the channels feed each other during processing.
  • Length dynamics: At fixed disclosure distance, influence declines as total context grows, supporting capacity pressure on transmission channels.Below-horizon totals leave influence alive and flat in the tested distance cells.
  • Mechanism conclusion: The gap is therefore a failure of integration rather than comprehension: the disclosure remains encoded and retrievable, while its transmission channels weaken with competing context.Both the fixed-size recurrent summary and attention-based lookup are causally implicated as substantial carriers.
  • Design implication: Workflow leverage lies in routing information to the point of judgment rather than only increasing computation after reading.The paper places design leverage between reading and judgment.

6 Workflow Architecture: Routing, Not Computation

The workflow experiments indicate that whether a disclosure influences judgment depends more on how information is routed than on additional computation. Generic summarization evicts decision-relevant content, while targeted structured restatement near the decision preserves or restores influence.

  • Decision-proximal routing: Targeted, structured restatement raises the disclosure’s long-context influence, with 67 percent surviving at 128k tokens versus 12 percent under baseline reading.All twelve firms respond in the same direction at 128k, and 128k influence exceeds the unassisted 2k baseline.
  • Workflow comparison: At 128k, extract-then-decide retains 67 percent of influence, whereas retrieve-summarize-decide reproduces none at any tested length.The contrast separates retaining the disclosure from routing it through budgeted generic summaries.
  • Generic summarization: Chunk-then-aggregate eliminates disclosure influence entirely: −0.001 at 128k and −0.001 at 2k.The target disclosure is absent from consolidated notes in twenty-four of twenty-four firm-arrangement cells at 2k and twenty-one of twenty-four at 128k.
  • Generic summarization: The failure occurs during extraction, as per-segment budgets fill notes with head-of-segment facts and evict the single-paragraph risk disclosure.The decision stage therefore receives no extracted risk content to use.
  • Routing, not computation: Added reasoning does not restore long-context influence; at short context, enabling reasoning reduces measured use from +5.5 to between +0.7 and +1.9 percentage points.In the tested vendor’s reasoning modes, more thinking about the same context appears to dilute short-context use.
  • Representation and authorship: A targeted extraction works similarly whether written by the model or an experimenter, while another model recovers about a quarter of the effect at 128k.The experimenter-written extraction is statistically indistinguishable from the model’s own; the other model’s interval includes zero.
  • Complete real filings: The workflow ordering replicates on complete statutory filings: whether a disclosure reaches recommendation is a property of workflow architecture, not the disclosure.The full-filing baseline leaves the disclosure without influence at −0.004, while the real-filing replication reproduces the gradient’s poles.
  • Representation and authorship: Passive source re-presentation recovers little, showing that the effective intervention is the representation’s targeted structure, not proximity alone.The restatement works while the filing remains in context, whereas replacing the document with notes removes influence altogether.

7 Conclusion

The paper concludes that retrieving a disclosure does not ensure its use in judgment: integration can fail after delegation, and workflow architecture determines whether information reaches decisions. Decision-proximal, structured representations restore influence more reliably than generic summarization, while model capability only shifts the integration horizon.

  • Core conclusion: A model that can recite a disclosure has not necessarily incorporated it into its judgment.The retrieval–integration gap is a post-delegation failure in which externally verifiable retrieval succeeds but the delegated judgment does not use the fact.
  • Core conclusion: Integration influence falls with context length even when retrieval remains accurate, separating machine readability from machine decision usefulness.The disclosure’s influence erodes as total length grows, while direct questioning can still succeed and retrieval-based checks may miss the divergence.
  • Mechanism: Causal memory interventions indicate that compressed summaries and source-text lookup jointly carry disclosures into judgments.Removing either memory channel reduces disclosure influence, while retaining either channel restores part of it; the estimated channel effects are statistically indistinguishable and not claimed to be additive.
  • Workflow architecture: Model capability and workflow architecture play distinct complementary roles: capability moves the integration horizon outward, while architecture determines whether information reaches judgment at all.Extended-reasoning modes do not restore lost influence, but a decision-proximal restatement raises full-length retained influence from 12 to 67 percent, including 8.5 percentage points at 128,000 tokens.
  • Implications: Evaluations should measure marginal decision influence rather than retrieval accuracy alone.Investors and intermediaries should test whether economically meaningful information changes produce appropriate judgment changes, because retrieval-based evaluations can certify invariant judgments.
  • Workflow architecture: Generic chunk-and-summarize workflows evict relevant disclosures, whereas targeted structured restatements adjacent to the decision restore their influence.Budgeted generic summaries evict a single-paragraph disclosure before decision time, while targeted extractions retain it; experimenter-written extractions work as well as model-generated ones.
  • Limitations: The study’s causal channel analysis is limited to one hybrid model family, although behavioral dissociations replicate across other families.How disclosures travel in architectures without a recurrent channel remains to be mapped, and real-filing evidence is described as capability-conditional rather than universal.

A.4 Real-Disclosure Redaction Sample

The redaction experiment removes real disclosures from actual 10-K filings to test whether those disclosures affect investment judgments.

  • Candidate disclosures came from 10-K filings filed February through August 2026, after every studied model’s training cutoff.
  • The sample spans six categories, including liquidity constraints, litigation exposures, concentration, and quantified regulatory liabilities.

B.1 Decision Readouts

The decision readout converts model outputs into a deterministic three-way investment choice and reports the sell probability.

  • The fixed prompt requires exactly one action code: buy, hold, or sell.
  • For open-weight models, sell probability is computed by softmax over the three final-position action-code logits.
  • Float32 recomputation of code logits addresses bf16 output-head quantization near 0.125.

B.2 Retrieval Readout

Retrieval is measured separately from judgment using directed fact extraction, frozen scoring keys, and pooled randomization tests under controlled runtime conditions.

  • Directed retrieval asks what the filing says about the target matter and scores answers against frozen per-firm fact keys.
  • Paraphrased amounts or statuses receive credit through alias groups without judgment calls during scoring.
  • The firm-filing is the inference unit, and 2,000 permutation draws form an empirical null from irrelevant pseudo-target insertions.
  • All paired conditions use a fixed library, kernel path, numeric precision, and device policy to pin runtime variation.
  • Results from a configuration affected by a multi-GPU communication defect were discarded and regenerated under a verified-correct configuration.

C Mechanism Methods

The mechanism analysis combines representation probes, sparse autoencoders, memory recombination, attention localization, and causal interventions to study how disclosures reach decisions.

  • Representation instruments: A vocabulary-direction probe measures risk-word rank at the disclosure’s fact position and immediately before the answer at the decision position.
  • Representation instruments: Probe validation across adjacent linear/full attention pairs finds readability tracks depth rather than attention type.
  • Representation instruments: Sparse autoencoders identify 24 risk features from activation contrasts and measure their summed activation at fact and decision positions.
  • Causal interventions: Three intervention designs use matched controls to test causal claims about decision-position risk subspaces, fact-position ablation, and memory transmission.
  • Causal interventions: Under float32 readout, all three intervention designs produce nulls with working controls, including t = −1.38 for ablation effects across firms.
  • Attention localization: Decision-time attention is measured by forwarding the final token against the preserved cache, while access interventions modify selected query-to-key weights.
  • Memory channels: The hybrid architecture maintains full-attention key–value stores and recurrent linear-attention states, which can be recombined across target and neutral runs.

E Robustness and Additional Results

Robustness checks show that severity-ordering ability is required for interpreting length decay, while changing filler placement does not produce a meaningful behavioral contrast.

  • Severity capability gate: A model unable to order disclosure severities at 2k tokens cannot support length-decay interpretation.The full-attention control fails this capability gate and is excluded from decay analysis.
  • Arrangement robustness: The before-minus-after filler-placement gap is statistically indistinguishable from the corresponding pseudo-target insertion gap at every tested length.The retention interpretation therefore rests on internal-representation evidence and the workflow gradient, not this behavioral contrast.

E.3 Sign Inversions

Long-context sign inversions are sparse and treated as suggestive rather than confirmed because some cells show disclosures reducing sell recommendations and chance alone predicts false positives.

  • Sign inversions: Sign inversions are reported as suggestive because approximately two false-positive stars are expected across forty tested cells.The reported pattern includes isolated mid-length inversions, a cross-family echo at adjacent lengths, and an arrangement-dependent flip in the post-trained 9B.

E.4 Repetition and Semantic Interference

Repeating a disclosure partially increases its influence, but the benefit is uneven and does not restore long-context use; semantically similar filler separates from unrelated filler only selectively.

  • Repetition: Fourfold repetition exceeds single-occurrence influence in pooled length cells, whereas twofold repetition fails to improve every cell.At 128k, four repetitions recover less than one third of a single mention’s 2k influence.
  • Semantic interference: Same-risk filler separates from unrelated filler only on the production model.

E.5 Out-of-Sample API Arm and Workflow Details

Additional experiments test the retrieval–use gap across models, tasks, interventions, and realistic filings, then identify workflow architecture as a key determinant of whether retrieved disclosures affect judgments.

  • Out-of-sample API arm: A dated API snapshot replicates flat directed-retrieval performance with zero neutral false positives, but its behavioral estimates are too noisy for firm-level claims.It is retained only as an out-of-sample retrieval witness.
  • Reasoning effort: Extended reasoning does not restore long-context use and reduces measured use at 2k for the production API model.
  • Workflow remedies: The extract-then-decide workflow retains only roughly 0.5–0.7 of its 2k use at 128k, showing residual erosion within the remedy.
  • Ecological validation: The retrieval–use dissociation replicates in complete investor-visible 10-K filings with inserted disclosures and removed conflicting passages.The model retrieves the inserted disclosure almost perfectly, but its investment judgment does not respond.
  • Workflow audit: Chunked notes routinely transcribe segment-head facts until budgets bind, evicting the single-point risk disclosure before the decision stage.The audit attributes zero influence to the extraction stage.
  • Primary contrast: The primary contrast measures each firm’s Use2k − Use128k difference, with the production API reported in percentage points of stated default probability.The production-API decline is not statistically resolved and should be read as a point-estimate pattern.
  • Second judgment task: The second economic judgment uses a one-token restricted-softmax probability readout, defining use as the target-minus-neutral contrast.
  • Position and distance: Disclosure influence is also examined by total context and disclosure-to-decision distance while holding the filler union byte-identical within rows.At 4,000 and 8,000 tokens, moving the disclosure from 1,000 to 5,500 tokens changes use by less than 0.0001 at fixed total length.
Loading 2608.24842v1…