Source-linked AI summary

Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline

Naimur Rahman

arXiv:2608.21559v1cs.CL

TL;DR

Multi-stage LLM pipelines may remain structurally valid when intermediate evidence deteriorates, creating a gap between parser validity and substantive reliability. This paper operationalises Evidence-State Reliability and evaluates it separately from parser validity in a controlled decision–audit–escalation experiment. Across matched degraded conditions, evidence-sensitive success deteriorated while parser validity moved positively, and detection did not produce recovery within the evaluated pipeline.

  • Problem

    Intermediate evidence can become incomplete, compressed, or conflicting while downstream outputs remain structurally valid, so parser validation alone cannot establish evidence sufficiency.

  • Method

    The study operationalises Evidence-State Reliability as a separate evaluation layer and uses controlled evidence interventions with matched uncertainty analysis across pipeline stages.

  • Results

    Across nine matched degraded-minus-clean comparisons, all operational stage-success estimates were negative with 95% bootstrap intervals below zero, while all parser-validity point estimates were positive.

  • Takeaways & Limitations

    Parser validity is necessary for automated pipeline operation but insufficient as a proxy for Evidence-State Reliability; audit detection and recovery must be evaluated separately.

  • Takeaways & Limitations

    The findings are confined to the study’s evaluated model configuration, pipeline design, selected sanitised cases, scoring procedure, and single scaled run.

Abstract

from arXiv · show

Multi-stage LLM pipelines can remain structurally valid even when evidence available to downstream stages becomes incomplete, compressed, or conflicting. This paper introduces and operationalizes Evidence-State Reliability (ESR), an evaluation layer concerned with whether intermediate evidence remains sufficiently complete, grounded, internally consistent, and usable for a stage's assigned function. ESR is evaluated separately from parser validity, which measures structural conformance. We evaluate the framework using GLM-5.2 on 60 sanitized base cases under four evidence conditions: clean, compressed-lossy, partial-dropout, and noisy-conflicting. Each condition was processed through decision, audit, and escalation stages. The design comprised 720 planned and ledgered calls, with 713 retained, sanitized execution rows. Across nine matched degraded-minus-clean condition-stage comparisons, all operational stage-success estimates were negative, and all 95% bootstrap intervals remained below zero. All nine parser-validity point estimates were positive, although the three partial-dropout intervals included zero. Among parser-valid degraded audit outputs, degradation detection was 1.0 in each degraded condition, while false-assurance rates remained non-zero; among parser-valid degraded escalation outputs, recovery was 0.0 in every degraded condition. The results show a bounded reliability-layer divergence in the evaluated pipeline: structural conformance can improve directionally while evidence-sensitive stage success deteriorates under the same controlled intervention. They also separate detection of degraded evidence from recovery. The conclusions are limited to the evaluated model configuration, pipeline design, selected sanitized cases, scoring procedure, and single scaled run.

1 Introduction

The paper introduces Evidence-State Reliability as a layer for evaluating whether intermediate evidence remains adequate for stage functions, separately from parser validity. A controlled decision–audit–escalation experiment tests whether degradation affects these measures differently and whether detection leads to recovery.

  • Parser validity measures structural conformance, whereas ESR concerns whether available evidence supports a stage’s substantive function.
  • Evidence-State Reliability evaluates whether intermediate evidence remains complete, grounded, internally consistent, and usable for a stage’s assigned function.
  • 60 sanitised base cases were represented under clean, compressed-lossy, partial-dropout, and noisy-conflicting evidence conditions across decision, audit, and escalation stages.Base-case identity was preserved for matched comparisons across conditions and stages.
  • Across nine matched degraded-minus-clean comparisons, all operational stage-success estimates were negative with 95% bootstrap intervals below zero, while all parser-validity point estimates were positive.The three partial-dropout parser-validity intervals included zero.
  • Audit detection did not imply recovery: false-assurance rates remained non-zero, and no parser-valid degraded escalation output was classified as recovered.The study treats detection and restoration of successful evidence-sensitive operation as separate pipeline functions.
  • The study contributes a stage-aware framework linking parser validity, evidence-sensitive success, audit detection, false assurance, escalation recovery, and sequence-level reliability patterns.Its conclusions are confined to the evaluated model configuration, pipeline design, controlled conditions, selected cases, scoring procedure, and single scaled run.

2 Related Work

Related work motivates multidimensional, component-level evaluation of structured outputs, evidence sufficiency, auditing, deferral, recovery, and cascading failure. This study integrates these concerns into a bounded, matched framework focused on intermediate evidence states and stage transitions.

  • Prior evaluation frameworks separate capabilities or properties such as faithfulness, context relevance, answer relevance, calibration, and task performance.
  • ESR extends multidimensional evaluation to intermediate pipeline evidence by asking whether a stage receives evidence capable of supporting its assigned function.
  • Structured-output research distinguishes formal validity from value correctness, executability, and task performance; this study compares structural validity with evidence-sensitive success under controlled changes.
  • The experiment processes controlled versions of the same base case through decision, audit, and escalation to assess degradation, detection, and downstream recovery.
  • The computational audit records degradation detection and false assurance, but it is not an independent institutional audit or a source of regulatory assurance.
  • The escalation stage measures recovery after degradation and is not a learned optimal deferral policy.
  • 2.4 Reliability cascades: Reliability cascade is used narrowly for diagnosed sequence patterns in which parser or evidence failures, false assurance, or incomplete execution prevent a preserved successful path.The taxonomy is not a general theory of cascading failure or an estimate of deployed-system failure prevalence.
  • The contribution is a bounded integration of these ideas within one controlled, matched multi-stage experiment rather than a claim that the individual ideas are new.

3 Evidence-State Reliability Framework

The framework treats Evidence-State Reliability as a stage-specific evaluation of whether intermediate evidence remains usable, grounded, complete, and internally consistent, separate from parser validity. It operationalizes this distinction through stage-success measures, matched comparisons, divergence metrics, and diagnostic cascade categories.

  • Evidence-State Reliability: Evidence-State Reliability evaluates whether evidence is complete, grounded, internally consistent, and usable for a receiving stage's assigned objective.
  • Operational measures: Parser validity measures structural conformance, whereas operational ESR requires parser validity plus a positive sanitised validity judgment.Rows lacking parser-valid output or a positive validity judgment are not counted as stage-successful.
  • Operational measures: Operational ESR is a stage-specific indicator rather than an exhaustive measure of the broader Evidence-State Reliability construct.
  • Reliability-layer divergence: Reliability-layer divergence measures the difference between structural-validity and evidence-sensitive success responses to the same intervention.A positive RLD indicates that parser validity changed more favourably than operational ESR.
  • Reliability-layer divergence: The study's target pattern is positive parser-validity change alongside negative ESR change, meaning the two evaluation layers provide opposing assessments.
  • Pipeline diagnostics: Audit detection, escalation recovery, and cascade taxonomy categories remain separate because recognising degradation does not establish successful remediation.The taxonomy distinguishes parser failure, detected-but-unrecovered degradation, audit false assurance, incomplete sequences, and preserved success.

4 Methodology

The methodology uses a paired, controlled within-case experiment that applies three defined evidence degradations against a clean baseline across 60 sanitised cases. It preserves case identity for matched comparisons while limiting inference to the selected experimental package and evaluated conditions.

  • Experimental design: 60 cases × 4 evidence conditions × 3 pipeline stages = 720 planned model calls.
  • Experimental design: Each degraded observation was compared with the corresponding clean observation for the same base case and pipeline stage.This paired design preserves base-case identity across conditions.
  • Scope: The selected cases are a controlled experimental set, not a representative or probability sample of consumer complaints or deployed AI interactions.
  • Evidence conditions: The clean condition is a non-degraded experimental baseline, not a guarantee of exhaustive evidence or parser-valid output.
  • Evidence conditions: The three degraded conditions model information loss through compression, information loss through omission, and evidential conflict through noise.
  • Scope: The conclusions exclude direct modeling of temporal staleness, distribution shift, adversarial manipulation, policy ambiguity, retrieval failure, and delayed ground truth.

4.5 Pipeline stages

The pipeline comprises decision, computational audit, and escalation stages, each evaluated with stage-specific computational indicators. Execution accounting and retained-row rules determine which outputs enter condition-stage and matched analyses.

  • Decision: The decision stage generates an assessment or recommendation and measures evidence adequacy and evidence-sensitive stage success.Its outputs are computational pipeline results, not real financial decisions or regulatory determinations.
  • Computational audit: The audit stage examines degradation or concern and evaluates degradation detection and audit false assurance among parser-valid audit rows.
  • Escalation: The escalation stage evaluates downstream response and records whether the implemented mechanism restored operational stage success among parser-valid escalation rows.
  • Stage transitions: Detection and recovery are separate functions: an audit can identify an evidential problem while escalation fails to restore successful evidence-sensitive operation.
  • Execution accounting: 470 parser-valid and 243 parser-invalid retained rows comprised the 713-row sanitised execution dataset.Every parser-valid ledger output was retained.
  • Analysis rules: Matched degraded-minus-clean comparisons included only cases with both observations retained at the same pipeline stage, producing paired sample sizes from 57 to 60.

4.7 Uncertainty analysis

The uncertainty analysis compares degraded-minus-clean measurements using matched-case bootstrap intervals and treats the results as uncertainty across the selected sanitized cases, not as population-level inference. The reporting artifacts support deterministic inspection of execution, condition-stage results, cascade classifications, figure data, and evidence mappings, but not exact replay of hosted-model interactions.

  • Three model-output-coded measurements were compared using matched degraded-minus-clean mean differences: evidence-sensitive stage success, parser validity, and evidence-state adequacy.
  • Matched base-case pairs were resampled together so bootstrap uncertainty preserved the clean–degraded relationship for each case.The procedure used a nonparametric bootstrap over matched base cases.
  • 2,000 bootstrap resamples and fixed random seed 5205 defined the uncertainty procedure.
  • 27 condition-stage-metric combinations received percentile-based 95% bootstrap intervals.
  • The intervals describe uncertainty across selected sanitized base cases under the implemented procedures, not population-level confidence for deployed systems or stakeholders.The analysis records whether intervals include zero without claiming statistical significance, population generalizability, or causal identification beyond the controlled within-case intervention design.
  • Committed artifacts enable deterministic verification of execution counts, condition-stage tables, cascade classifications, figure source data, and claim-to-evidence mappings.They provide artifact-level reproducibility and auditability, but raw prompts, responses, records, archives, credentials, and environment files are excluded, preventing exact prompt-response replay.

5 Results

Across the evaluated conditions, parser validity improved directionally while evidence-sensitive stage success deteriorated, and downstream detection did not yield escalation recovery.

  • Execution and parser accounting: 713 retained sanitised execution rows came from 720 planned model calls.Missing observations occurred only in clean and noisy-conflicting conditions, leaving retained condition-stage denominators from 58 to 60.
  • Condition-by-stage outcomes: Every degraded condition-stage cell had higher parser-validity than its clean counterpart, while every degraded cell had operational stage-success of 0.0.The clean parser-validity baseline rates were 0.508475 at decision, 0.517241 at audit, and 0.416667 at escalation.
  • Paired bootstrap results: All nine parser-validity point estimates were positive, but the three partial-dropout intervals included zero.The partial-dropout escalation interval was [−0.066667, 0.300000].
  • Paired bootstrap results: All nine operational stage-success estimates were negative, with every corresponding 95% bootstrap interval below zero.Parser validity and operational stage success therefore moved in opposite directions across every evaluated degraded condition-stage comparison.
  • Audit detection and escalation recovery: Escalation recovery was 0.0 in every degraded condition among parser-valid escalation outputs.Recovery was 0 of 43 for compressed-lossy, 0 of 32 for partial-dropout, and 0 of 37 for noisy-conflicting evidence; the clean parser-valid escalation success rate was 1.0.
  • Sequence-level reliability cascades: Across 240 condition-linked sequence groups, 223 were classified as operational cascade failures and 17 as preserved successes.The aggregate rate was 223/240 = 0.929167 and includes both clean and degraded conditions.
  • Sequence-level failure composition: The sequence taxonomy comprised 143 parser-failure cascades, 71 detected-but-unrecovered patterns, three audit-false-assurance patterns, six incomplete sequences, and 17 preserved successes.These mutually exclusive categories sum to 240 groups.

6 Discussion

The discussion frames the findings as a divergence between structural conformance and evidence-sensitive reliability under controlled degradation. It also distinguishes degradation detection from recovery and supports layered evaluation rather than parser validation alone.

  • Reliability-layer divergence: All nine parser-validity point estimates were positive, while all nine operational stage-success estimates were negative.The three partial-dropout parser-validity intervals included zero, whereas all stage-success intervals remained below zero.
  • Reliability-layer divergence: Parser validity and operational ESR measured different properties rather than interchangeable notions of reliability.Schema compliance can appear structurally improved while evidence-sensitive performance deteriorates.
  • Implications: Parser validity remains necessary, but it should be evaluated alongside evidence-sensitive measures rather than treated as a sufficient reliability proxy.The proposed layered approach also separates detection, false assurance, recovery, and sequence completion.
  • Measurement interpretation: The study treats operational stage success as an ESR indicator, not as an exhaustive measurement of Evidence-State Reliability.Stage success required both parser-valid output and a positive model-output-coded validity judgment.
  • Measurement interpretation: The separately recorded evidence-state adequacy variable also had all nine degraded-minus-clean intervals below zero.Both indicators were derived from model-output-coded fields rather than independent human or expert adjudication.
  • Detection versus recovery: Detection reached 1.0 in every degraded audit condition, but false-assurance rates remained non-zero and escalation recovery was 0.0.The implemented audit indicator detected degradation without restoring successful evidence-sensitive operation.
  • Implications: The results support rules governing evidence sufficiency, conflict or absence representation, deferral, and post-detection changes in pipeline design.Structural contracts specify machine-readable format but not the evidential conditions required for substantive functions.

7 Threats to Validity and Limitations

The validity limits concern construct measurement, implementation dependence, sampling and execution scope, intervention coverage, and reproducibility. These constraints bound interpretation to the evaluated design and selected cases rather than deployed systems generally.

  • Construct and measurement validity: Operational stage success and related indicators are model-output-coded proxies, not exhaustive or independently adjudicated measurements of ESR.Independent assessments would be needed to evaluate evidence sufficiency, grounding, consistency, audit correctness, deferral, and recovery.
  • Internal validity: The observed divergence may depend on parser-contract difficulty, stage-specific instructions, and interactions among these implementation elements.The paired design controls for base-case identity but not all properties of the experimental implementation.
  • Baseline interpretation: Clean parser-validity rates ranged from approximately 0.42 to 0.52 across stages, creating headroom for positive degraded-condition changes.The results therefore do not establish that degradation intrinsically improves structural validity.
  • External validity: The study used one GLM-5.2 configuration, one pipeline, one scaled run, 60 selected cases, one scoring contract, and three degradation families.Its findings do not establish cross-model, provider-independent, domain-independent, or universal behavior.
  • External validity: The 60 cases are not claimed to represent a probability sample of complaints, financial decisions, or deployed AI interactions.Bootstrap intervals describe variability across selected matched cases, not population sampling uncertainty.
  • Intervention coverage: The interventions model compression, omission, and conflict, but not adversarial manipulation, malicious injection, or changing model or provider behavior.Replication across models, domains, output contracts, evidence sources, and recovery architectures is required for broader generalization.
  • Conclusion validity: All nine stage-success intervals remained below zero, but three partial-dropout parser-validity intervals included zero.The evidence supports consistent negative stage-success changes without uniformly interval-separated parser-validity improvement.
  • Conclusion validity: The aggregate cascade-failure rate of 0.929167 is not a pure degradation effect, deployment failure probability, consumer-risk estimate, or prevalence estimate.It includes clean-condition failures; paired condition-stage comparisons are the stronger basis for interpreting degradation-associated changes.

8 Ethical and Data-Use Considerations

The study uses sanitized CFPB-derived evidence solely to evaluate a controlled pipeline mechanism, not to make claims about people, organizations, complaints, or deployed financial systems.

  • Data use: Sanitized derivative CFPB evidence packets were used solely to evaluate a controlled reliability mechanism within a multi-stage language-model pipeline.The materials were constructed from public complaint-context data.
  • Use boundaries: The experiment did not evaluate live consumer decisions or produce actual credit, lending, regulatory, legal, or financial determinations.No output affected an individual, organization, account, complaint, or institutional process.
  • Interpretive boundaries: CFPB complaints were not treated as adjudicated findings or a representative sample of consumer experience.The study does not infer complaint truth, misconduct, violations, consumer harm, decision appropriateness, or deployment frequency.
  • Data retention: Only sanitized, derived, and aggregate research artifacts were retained within the reported evidence package.Direct personal identifiers and JSONL execution archives were among the excluded materials.
  • Interpretive boundaries: The computational outputs are not professional financial advice, institutional auditing, regulatory assurance, or human expert adjudication.Findings should be interpreted strictly as evidence about the evaluated pipeline under controlled interventions.

9 Reproducibility and Availability

The research package provides artifact-level traceability and independent verification of the reported analysis, while excluding materials needed for unrestricted reconstruction of every hosted-model interaction.

  • Available artifacts: The repository contains metric definitions, condition-stage estimates, matched comparisons, bootstrap intervals, audit and escalation results, cascade classifications, and claim-to-artifact mappings.These materials support verification of the reported analysis components.
  • Repository checkpoints: The principal empirical analysis was generated from repository checkpoint c3e802c71976faae34ac3f327b537e12916bc970.A later checkpoint, 3144d46e3c59b011f2939d7ab3a590c688543492, was used for manuscript preparation and submission validation.
  • Verification scope: The shared artifacts support deterministic verification of planned, ledgered, retained, parser-valid, and parser-invalid execution accounting.They also include metric formulas and reported condition-stage outputs.
  • Reproducibility boundary: Raw records, prompts, responses, execution archives, credentials, and environment files are excluded from the repository.The package therefore does not support exact prompt-response replay or complete computational replication.

10 Conclusion

The study operationalises ESR alongside parser validity and finds a consistent divergence: evidence-sensitive stage success deteriorated under controlled degradation even as structural conformance moved directionally upward. Audit detection did not ensure recovery, and the conclusions remain bounded by the evaluated setup.

  • Framework: ESR was operationalised through structural parser validity, evidence-sensitive stage success, controlled interventions, uncertainty analysis, audit outcomes, escalation recovery, and sequence-level patterns.The framework treats these as complementary measurements of multi-stage pipeline reliability.
  • Scope: The evaluation used one GLM-5.2 configuration, one pipeline design, one scaled run, 60 selected sanitised cases, three degradation families, and model-output-coded scoring.Cross-model, multi-run, multi-domain, and independently adjudicated replication is required before broader claims of generality or deployment reliability.
  • Empirical findings: All nine matched degraded-minus-clean comparisons had negative operational stage-success estimates, with every 95% bootstrap interval below zero.All nine parser-validity point estimates were positive, although three partial-dropout intervals included zero.
  • Empirical findings: Degradation detection was 1.0 in every degraded audit condition, while false-assurance rates remained non-zero and escalation recovery was 0.0 throughout.These results distinguish identifying degraded evidence from restoring successful operation.
  • Empirical findings: 223 of 240 condition-linked groups met the operational cascade-failure definition, spanning parser failure, detected-but-unrecovered degradation, false assurance, incomplete execution, and preserved success.The taxonomy distinguishes failure families that imply different technical responses.
  • Implications: The findings support parser validity as necessary for automated operation but insufficient as a proxy for Evidence-State Reliability.Structural conformance does not establish that evidence remains adequate for a stage’s substantive function.
Loading 2608.21559v1…