Source-linked AI summary

Selection-Aware Stress Testing for Interactive Agents

Yang Xu, Chenang Li, Jiefu Zhang, Haixiang Sun, Zhou Li, Vaneet Aggarwal

arXiv:2608.30916v1cs.LGstat.APstat.ML

TL;DR

Interactive-agent evaluations can select both a workflow and its apparent weak task regime from the same benchmark, risking non-replicating reversals. SASST learns a task reweighting from discovery tasks and tests the frozen comparison on separate confirmation tasks with joint uncertainty bounds. In the agent studies, a +3.75 pp discovery gain vanished on confirmation, and neither model produced a confirmed workflow benefit or stable stress rule.

  • Problem

    Benchmark-based workflow selection and post hoc stress-task search can reuse the same outcomes, making apparent reversals require independent confirmation.

  • Method

    SASST learns bounded task reweighting from pre-execution features on discovery clusters, checks support and stability, freezes the rule, and evaluates the same comparison on separate confirmation clusters.

  • Results

    +3.75 pp Qwen3 discovery gain vanished on confirmation, while neither model yielded a confirmed workflow benefit or reusable stress description.

  • Takeaways & Limitations

    The protected report supports neither a workflow-benefit claim nor a stable stress-rule claim when discovery findings fail confirmation or stability checks.

  • Takeaways & Limitations

    The theorem assumes conditionally i.i.d. confirmation clusters, while the studies use forty confirmation tasks and a second model on the same τ-bench task pool.

Abstract

from arXiv · show

Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks. The protocol checks support and stability, uses joint bounds for all planned claims, and can return no claim. We prove conditional asymptotic validity under stated cluster assumptions. A forty-cluster audit finds Gaussian undercoverage and conservative Bonferroni $t$ bounds. In one 480-episode $τ$-bench study, a $3.75$ point discovery gain vanished on confirmation. A second-model study likewise confirmed neither a workflow benefit nor a stable stress rule.

1 Introduction

SASST addresses the risk that workflow selection and stress-task discovery reuse the same benchmark evidence. It separates discovery from confirmation while testing task emphasis rather than changing workflows or the comparison.

  • A workflow that wins on average may lose its advantage when a coherent task type receives more emphasis.The apparent reversal may reflect sampling noise when the subgroup is selected after inspecting the same outcomes.
  • SASST learns a reusable task-weighting rule from pre-execution task attributes and rejects rules with insufficient support or excessive shift.
  • SASST freezes the selected rule, retests the same comparison on separate confirmation tasks, and uses joint uncertainty bounds for planned claims.Its report distinguishes confirmed, not confirmed, failed discovery, failed feasibility, and abstained outcomes.

2 Method

SASST fixes the workflow comparison, learns a bounded task reweighting from discovery data, and evaluates that frozen rule on independent confirmation clusters. Its validity result is conditional on explicit cluster and regularity assumptions.

  • Fixed comparison and stress target: SASST keeps workflows and the comparison plan fixed while changing only task emphasis through weighted workflow contrasts.For example, the target can compare Plan + Verifier with ReAct under identical task weights.
  • Fixed comparison and stress target: A stress rule uses bounded nonnegative scores from pre-execution features to reweight represented task types within allowed distribution shifts.The stress target does not cover absent tasks or arbitrary deployment traffic.
  • Discovery and confirmation: Discovery ranks candidate rules, freezes the selected rule before confirmation, and confirmation recomputes the exact paired contrast rather than reusing the ranking objective.
  • Discovery and confirmation: Density, effective-sample-size, radius, and stability checks limit concentration, distributional shift, and unstable rule descriptions.Confirmation repeats support checks and forms a joint uncertainty bound over all planned claims.
  • Validity and scope: Under conditionally i.i.d. confirmation clusters and stated regularity conditions, Gaussian multiplier inference consistently approximates the joint law and provides conditional coverage.The protected objects are candidate contrasts rather than the argmax under hard workflow selection.
  • Validity and scope: The theorem does not directly justify the finite-sample Bonferroni t safeguard, and the stratified designs in Studies F and G require an extension not supplied by the theorem.

3 Validation

Validation tests conditional coverage, small-cluster behavior, and detection of supported stress signals. Results show separated discovery and confirmation improve coverage, while conservative safeguards help in the forty-cluster regime.

  • Coverage validation: 94.75–95.30% pointwise and 94.95–95.50% six-slot simultaneous coverage were observed for M ≥1000.
  • Coverage validation: At M = 1000, same-data worst-tercile coverage was 91.6%, versus 93.6% when discovery and confirmation were separated (p = 0.032).
  • Small-cluster safeguard: At N = 40, Gaussian coverage ranged from 90.55–93.50%, while Bonferroni-t coverage ranged from 96.40–97.55% across four null settings.An initial audit reported 92.2–92.8% Gaussian coverage.
  • Signal detection: With 200 clusters and adequate stressed ESS, the planted-fragility control selected the target in 92.1% of trials and detected it end-to-end in 64.6%.Conditional confirmation was 70.1%, with 1.7–1.8% null familywise error.
  • Rule-search safeguards: Rule-search audits found that geometry, declared-radius, support, and stability safeguards can exclude dominated, shifted, concentrated, or unstable rules.

4 Agent studies

Two-model τ-bench studies apply SASST to workflow benefits and task shifts. The discovery advantage did not survive confirmation, and neither model produced a reusable stress description or confirmed workflow benefit.

  • Study design: Each model study used 80 τ-bench tasks, two simulator seeds, three workflows, and 480 episodes, splitting whole tasks 40/40 for discovery and confirmation.Rules below the predeclared 0.60 stability threshold were abstained.
  • Results: The Qwen3 discovery advantage disappeared on confirmation, while Qwen2.5 had lower overall success, selected a different workflow, and produced different stress rules.
  • Results: Maximum rule stability was only 0.37 for Qwen3 and 0.27 for Qwen2.5, so neither model yielded a reusable stress description or confirmed workflow benefit.
  • Scope and limitations: The studies do not establish equivalence or safety because they have forty confirmation clusters, stressed ESS near 20, many universally failing tasks, and a reused task pool.The Qwen3 kernel bandwidth also made most task embeddings nearly orthogonal.

5 Decision interpretation and scope

SASST reports a decision status for each predeclared claim, distinguishing confirmation from non-confirmation, infeasibility, instability, and abstention. Its conclusions are limited to the benchmark and allowed task shifts, not deployment security.

  • Decision interpretation: The Qwen3 study confirmed neither the workflow benefit nor a stable stress rule after its discovery gain vanished on confirmation.The study reports a +3.75 percentage-point discovery gain, but the average gain disappeared on confirmation and the segment description was unstable.
  • Decision interpretation: SASST reports the frozen rule, estimated comparison, uncertainty bound, support, stability, geometry diagnostics, and final status for every planned claim.These fields make missing claims interpretable through support, stability, or uncertainty failures.
  • Scope and limitations: The theorem requires conditionally i.i.d. confirmation clusters, bounded task weights, stable denominators, honest discovery, and smooth comparison functions.The Bonferroni t safeguard is supported by four audited null settings rather than a universal finite-sample theorem.
  • Scope and limitations: Fresh task clusters, score-blind geometry calibration, policy-aligned outcomes, independent checking where relevant, and stressed-ESS power planning are required for the next protected study.More seeds on the same tasks do not replace independent clusters.
  • Scope and limitations: SASST supports decisions only for the benchmark and allowed task shifts, not as a deployment-security certificate.It distinguishes supported, underpowered, out-of-scope, and unstable claims before release.
  • Relation to prior evaluation: Existing agent benchmarks make workflow evaluation concrete but do not provide confidence bounds after task mixtures are selected through workflow search.Related methods study weak groups, subpopulation shifts, simultaneous inference, held-out replication, and robust optimization, whereas SASST audits selection-aware task mixtures.

B Proofs and technical details

The technical development linearizes weighted-ratio reports under cluster assumptions and tracks both held-out evaluation and splitwise selection effects. Taylor expansion and finite-dimensional arguments establish the stated plug-in approximation uniformly over the finite family.

  • Weighted-ratio expansion: The weighted-ratio expansion centers the empirical tilted mean over a common denominator and uses Slutsky’s theorem to obtain the approximation.The vector extension holds coordinatewise when the dimension is fixed.
  • Weighted-ratio expansion: The same weighted-ratio statement extends coordinatewise to a fixed-dimensional vector.This supports applying the expansion to the finite collection of report components.
  • Cluster linearization: Cluster linearization defines Φ_k,m from the gradient of F_m at the conditional mean and the cluster-level deviation H_k,m − µ_m.The lemma is conditional on discovery data and the assumptions of Theorem 1.
  • Proof conditions: Positive denominators, cluster laws of large numbers, finite report families, finite second moments, and cluster Lindeberg conditions control the Taylor remainder and empirical mean rate.The proof obtains O_p(N^-1/2) mean deviations and o_p(1) remainder terms uniformly over the finite family.
  • Cluster linearization: The coordinates of H_k,m include weighted-scoring and held-out numerators and denominators, plus selector and aggregation inputs.Thus the linearized contribution covers both report construction and adaptive selection inputs.
  • Cluster linearization: Held-out evaluation and splitwise selection are the two channels through which an evaluation unit changes an adaptive report.The resulting influence contribution contains both effects.

B.3 Proof of Theorem 1

The protocol conditions inference on discovery, establishes a joint Gaussian multiplier approximation under cluster assumptions, and protects simultaneous and post-selection reports. Algorithmic safeguards freeze the rule before confirmation, check support and stability, and distinguish confirmed, unsupported, infeasible, and abstained outcomes.

  • Proof of Theorem 1: Conditioning on discovery fixes the discovered rules and their random targets for confirmation inference.
  • Proof of Theorem 1: The centered estimator converges jointly to a Gaussian limit through a scalar CLT, Cramér–Wold, and Slutsky’s theorem.
  • Proof of Theorem 1: Conditional multiplier draws are Gaussian with covariance bΣ, yielding conditional weak convergence and pointwise bootstrap coverage.
  • Proof of Theorem 1: The simultaneous event supports every coordinate, including one selected after inspecting estimates and bands.
  • Executable protocol: The algorithm freezes encoders, candidate families, thresholds, workflows, and inference choices before confirmation outcomes are read.
  • Executable protocol: Discovery selects feasible rules using pre-execution features, while confirmation reapplies the frozen rule and checks support, stability, exact paired contrasts, and joint bounds.

D.4 Study B and the geometry audit

Study B compares alternative geometries for recovering stress rules, while the geometry audit tests whether a common radius can support the intended task shifts. The studies show geometry-dependent recovery and explicit feasibility limits.

  • Study B: 2.125% of KMS selections and 0.625% of max-sliced selections changed after correcting the Wasserstein routine, without changing the qualitative result.The correction affected selection frequencies but not the study’s qualitative conclusion.
  • Study B: MMD was strongest on three of four mechanisms, while max-sliced won on policy by tool; no geometry won every mechanism.These results support nonlinear rule recovery for some mechanisms but not universal KMS superiority.
  • Geometry audit: The target rule fell outside the common 75% radius budget for every geometry.The audit therefore identifies a support failure under the shared radius constraint.
  • Feasibility and abstention: Study C’s support gates blocked infeasible or unstable rules across five synthetic settings, while Study D abstained when its detectable effect had unstable functional description.The safeguards cover density, ESS, split denominators, and stability rather than forcing a claim.
  • Controller audit: All twenty mechanical gates and ten boundary and regression tests passed, but the audit did not validate informed consent or model-based semantic verification.The audit applies to the revised f2_v2 controller, whereas Studies F and G used the historical study_f_v1 controller.

D.8 Studies F and G: interactive agent evaluations

Studies F and G apply the same discovery-confirmation protocol to two models on the same τ-bench task pool. The discovery advantage did not survive confirmation, and neither model produced a confirmed workflow benefit or reusable stress rule.

  • Experimental design: 480 episodes were collected in each study using 80 task IDs, two simulator seeds, three workflows, and a 40/40 task-cluster split.Study G is an across-model replication of the reporting decision on the same tasks, not an independent benchmark replication.
  • Support checks: Study F’s maximum density ratio was 2.73, within κ = 5, whereas Study G’s third radius failed feasibility at ESS = 18.7 below the minimum of 20.The two studies therefore differed in feasibility for one prespecified radius.
  • Results: Neither model yielded a reusable stress description or confirmed workflow benefit.Study G produced zero confirmed claims despite selecting different rules from Study F.
  • Sensitivity: Treating three parser errors as missing left both confirmation contrasts and all statuses unchanged; stronger balanced deletion changed two slots to failed feasibility while retaining zero confirmed claims.The sensitivity analysis changes status labels for some slots without changing the study-level absence of confirmed claims.
  • Post-hoc decomposition: In Study G, Plan + verifier completed 76/160 episodes versus 118/160 for ReAct, but completed-episode success was 23.7% versus 16.1% and was post-treatment selected.The decomposition should not be interpreted as a causal quality improvement.

D.9 Reproducibility and superseded analyses

The paper restricts its reported results to corrected artifacts and audits their arithmetic, provenance, claim wording, and exclusion of superseded analyses. A prospective natural follow-up is left for later work.

  • Artifact control: The paper tables use only corrected artifacts, and an arithmetic audit verifies discovery gains, confirmation contrasts, protected bounds, critical values, and status counts.The claim source matrix links each claim to a corrected artifact and records allowed and forbidden wording.
  • Artifact control: An exclusion audit checks that no paper script reads a superseded first-pass result.This audit complements the provenance and claim-source records.
  • Future work: A prospective natural follow-up would require fresh tasks, a valid policy-aligned outcome, predeclared inference, corrected feature geometry, a nonredundant intervention, and a stressed-ESS power plan.The paper leaves that experiment to later work.

E Limitations and broader impact

The paper’s guarantees and agent evidence are bounded by asymptotic assumptions, finite-cluster calibration, the declared stress-rule registry, and a narrow evaluation setting. It frames SASST as evaluation and release control, not deployment-safety proof.

  • Statistical scope: The theorem assumes bounded density ratios, stable denominators, honest discovery, and plausibly independent clusters.The forty-cluster audit found Gaussian undercoverage and conservative Bonferroni t bounds; the safeguard was audited in four null settings rather than proved for every finite sample.
  • Search scope: Stress-rule search is limited by the declared registry, feature map, radius, and support thresholds; rules outside the radius are outside audit scope.Poorly scaled high-dimensional kernel bandwidths can also make features nearly orthogonal.
  • Agent-study scope: The agent studies use one benchmark task pool, low-success open models, the same model as actor and user simulator, and forty confirmation clusters.The second-model study uses the same tasks and is therefore a model replication rather than an independent benchmark replication.
  • Broader impact: SASST is intended for evaluation and release control, distinguishing supported workflow claims from development-set gains, unstable segments, or low-support analyses.The paper rejects treating benchmark relative results as proof of deployment safety or non-confirmation as equivalence.
Loading 2608.30916v1…