Source-linked AI summary

Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable

Pranav Aggarwal

arXiv:2608.27167v1cs.AIcs.CL

TL;DR

This paper asks whether authoritative evidence genuinely informs LLM agents or merely licenses action on unknowable questions, testing the distinction by intervening on presented evidence. It finds that fabricated panels trigger commitment much like genuine data, while the act/don’t-act gate can be trained but remains fragile to prompt format.

  • Problem

    Prior work showed that relevant-looking context makes agents act on irreducibly uncertain questions, but whether information or authoritative presentation causes this remained unresolved.

  • Method

    The paper intervenes on evidence presentation, evaluates matched unknowable and answerable questions, and tests gate training across domains and response formats.

  • Results

    Fabricated panels induce commitment like genuine data, while supervised fine-tuning suppresses commitment and preserves answerable performance but fails under rigid formats.

  • Takeaways & Limitations

    The failure lies in the act/don’t-act gate rather than belief or judgment, making it trainable but not yet a robust safety property.

  • Takeaways & Limitations

    The observed effect comes from one developer’s models, with only three of twelve carrying it, so pooled rates do not generalize to frontier models broadly.

Abstract

from arXiv · show

An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.

1 Introduction

On provably aleatoric questions, authoritative-looking evidence makes LLM agents act despite unknowability, even when the evidence is fabricated; the failure lies in the action gate, which is trainable but context-fragile.

  • Core finding: 6.5% to 54.0%: adding an authoritative-looking indicator panel increases commitment across 12 frontier models, while fabricated panels produce comparable commitment to real data.Fabricating six indicators leaves commitment at 37.6% to 38.3%, and fabricating the entire panel yields 36.8% versus a 24.5% no-panel baseline.
  • Scope: The study targets aleatoric questions whose answers will resolve but are unknowable in advance, distinct from epistemically unanswerable queries involving missing information or flawed premises.Examples include short-horizon asset direction, match outcomes, and weather ten days ahead.
  • Failure localization: The failure is concentrated at the action gate: commitment varies from 0-100% across models and domains without tracking capability, while belief barely moves and matched answerable questions remain essentially perfectly solved.Stated probabilities score worse than a climatological baseline, so standard calibration metrics do not reveal the action failure.
  • Training intervention: Synthetic-only supervised fine-tuning drives commitment to 0.0% on original cases and transfers to three unseen domains across six independent runs.This supports trainability, while the intervention’s scope is qualified by its dependence on the response format.
  • Failure boundary: The trained gate holds when responses leave room for reasoning but fails under rigid formats that remove that room, making deployment context-fragile.The introduction reports reasoning in 240/240 and 288/288 responses when a slot exists, versus 0/288 when it does not.

2 Method

The method tests whether agents distinguish unknowable future outcomes from answerable questions under matched evidence panels, using domain-specific horizons, manipulated displays, and answerable controls. It defines commitment, discrimination, calibration, uncertainty, and denominator conventions while documenting important confounds and limitations.

  • Unknowability oracle: The study treats equity, crypto, sports, and weather forecasts at stated future horizons as unknowable, with crypto serving as the primary transfer domain and weather interpreted cautiously.Crypto has sealed outcomes and a verified near-chance base rate; weather lacks outcome checks and is therefore the weakest instrument.
  • Evidence gradient: Matched evidence conditions manipulate panel content, while scrambled arms replace technical fields with earlier values and can fabricate the header and regime tag.Transfer-domain L2 is byte-identical to the equity study’s thin arm, whereas the richer equity panel adds EMA20/50, MACD, ATR, and regime information.
  • Matched answerable controls: Every transfer domain pairs unknowable items with 36 YES and 36 NO answerable threshold questions recomputed from the same rich panels, preventing refusal-everything strategies from scoring perfectly.The answerable arm tests whether models can resolve questions from displayed data rather than merely decline future-tense prompts.
  • Agentic protocol: The agentic protocol presents a present anchor, a web_search tool limited to information available today, and a choice among ANSWER, CALL_TOOL, and DECLINE without a humility anchor.Pilot wording that supplied a safe answer, such as “50 = coin flip” or an explicit UNKNOWABLE option, suppressed the effect.
  • Metrics and analysis: Commitment is choosing ANSWER, discrimination is P(decline | unknowable) − P(decline | answerable), and evaluation uses sealed-outcome Brier scores, CORP calibration, case-clustered intervals, and fixed denominator rules.Unreadable responses remain non-commitments, failed API calls are dropped, and all comparisons use the same commitment definition.

3 The display, not the data, is the trigger

Professional-looking panels trigger commitment even when their contents are fabricated: fully invented displays produce nearly the same commitment as real data, and commitment rises with display density. The effect is concentrated in a few models, revealing sensitivity to authoritative packaging rather than truth.

  • The display, not the data, is the trigger: Scrambled indicators retain the real format and asset while substituting earlier-date values, preserving plausible headers and avoiding look-ahead or detectable inconsistency.This isolates whether models respond to authoritative presentation when the displayed technical values are informationally empty.
  • The display, not the data, is the trigger: Among the three models carrying the effect, fabricated minus real commitment is +11.11pp, with a 90% CI of [+4.17, +18.06].The pooled contrast masks heterogeneity: five models never vary in commitment, while three are more seduced by fabricated than real panels.
  • The display, not the data, is the trigger: Commitment rises from 0.0% with 0 indicators to 5.4%, 32.1%, and 50.0% with 2, 4, and 7 indicators, respectively.The models’ elicited edge rises alongside display density, from 2.0 to 3.1 to 9.2 to 14.1.
  • The display, not the data, is the trigger: Reasoning traces show models citing wrong-date indicators as evidence, while the same models describe fair-coin and other questions as genuinely unpredictable when no authoritative display drives commitment.The trace evidence makes the failure legible as reliance on packaging rather than valid information.
  • The display, not the data, is the trigger: No panel yields 24.5% commitment, versus 37.6% with real data, 38.3% with fabricated indicators, and 36.8% with the entire panel fabricated.Adding pure invention changes commitment nearly as much as adding factual data, showing that the display—not its truth value—is the trigger.

4 Evidence-induced commitment occurs across domains

Professional evidence sharply increases commitment on unknowable questions, peaking at 54.0% despite performance worse than an uninformative baseline. The effect transfers across crypto, sports, and weather and varies independently of model capability.

  • Confirmatory gradient: 54.0% commitment under the full panel versus 6.5% for the bare question, a +48pp increase, while a different entity’s panel reduces commitment to 3.5%.The evidence gradient also passes through 14.8% when models see two prices.
  • Cross-domain transfer: 34.7% commitment in crypto, 36.8% in sports, and 55.7% in weather at the fullest evidence level show the effect across domains.The corresponding bare-to-full ranges are 9.4% to 34.7%, 4.2% to 36.8%, and 9.1% to 55.7%; the gradient is monotone in two domains.
  • Model heterogeneity: Per-model commitment spans the full range and does not track capability, with Haiku at 0.0% and Opus at 86.1% on transfer items.The same comparison also reverses across experiments: Haiku commits on 45.8% of §3 crypto panels, whereas Opus commits on 12.5%.

5 The judgment is present; the gate does not consult it

Models can recognize that questions are unknowable, yet their action policy often ignores that judgment and commits anyway. The failure is a separable act/decline gate, not incapacity or substantial belief updating.

  • The judgment is present; the gate does not consult it: 90% of classifications labeled the question irreducible, yet only 0.4% of those cases triggered commitment, locating the failure in the act/decline gate.At the full-panel level, conditional commitment was 0.9% (4 of 910 pooled over bare and full-panel conditions).
  • The judgment is present; the gate does not consult it: A triage instruction cuts commitment from 54.0% to 10.2%, whereas matched-length diligence instruction achieves 47.6%, indicating that targeted guidance—not generic prompt pressure—can engage the gate.The triage reduction has 95% CI [-47, -41], while diligence raises commitment for some models.
  • The judgment is present; the gate does not consult it: As action changes by 48 points, mean |stated probability - 50| moves only from 4.7 to 7.7, while Brier Skill Score is -0.124 against a climatological forecast.Expected calibration error varies from 0.079 at 5 bins to 0.211 at 10 and 0.208 at 15, making binning-sensitive screening unreliable here.
  • The judgment is present; the gate does not consult it: 98.6–100% of matched answerable controls were answered with 98.6–100% per-model accuracy and 99.9% pooled accuracy, showing the models can use the panels when questions are resolvable.There was a single error on one sports item among 859 answered controls.
  • The judgment is present; the gate does not consult it: AUROC 0.887 and 83.7% best-threshold accuracy for the elicited edge signal exceed the models’ 67.9% act/decline accuracy, showing the signal is separable from action.The comparison is between different quantities—AUROC and accuracy—and does not treat them as directly differenced.

6 The gate can be trained

Supervised fine-tuning on 540 synthetic cases trains a separable act/don’t-act gate that transfers across crypto, sports, and weather, while preserving answers to knowable questions. The gate is robust to rephrasing but fragile under structurally novel prompts, especially when refusal requires distinguishing knowability from grammatical tense.

  • Training recipe: Training uses completion-only supervised fine-tuning with chain-of-thought targets, 4-bit QLoRA, three epochs, and 540 synthetic cases dominated by dice, coins, jars, timers, and calendars.The training set contains no stocks, crypto, or weather items, although it includes 24 sports items sharing a transfer template.
  • Natural-framing transfer: Youden’s J reaches +62 to +100 for crypto, +83 to +100 for sports, and +75 to +100 for weather across six independent runs under natural framing.The preference continuation scores +100 / +100 / +88 and falls within every independent-run range.
  • Template-overlap control: Removing the 24 sports cases leaves sports Youden’s J at +83 to +100 across four independent runs, matching runs that included them, while widening variation elsewhere.The result supports transfer rather than simple recall, though the 516-case recipe spans +46 to +100 on crypto versus +62 to +100 for the 540-case recipe.
  • Knowability control: On a tense-balanced control, the model answers 95.8% of answerable-future items and declines every unknowable-present item, showing discrimination beyond a future-tense refusal rule.An independently trained run answers 100% of answerable-future items, while the preference-trained variant answers 79.2%.
  • Structural fragility: Under structurally novel frontier framing, within-tense J falls to +18 to +19pp for present items and +62 to +79pp for future items, with only 20-28% of unknowable-present items declined.Answerable-future items remain comparatively intact, with 0-17% wrongly declined; explicit randomness words remain a possible residual cue.

7 The trained gate transfers to the original cases

On the original 40 cases, the trained gate transfers by reducing commitment to zero at every level while preserving perfect performance on matched answerable questions. The refusal behavior improves further with preference training, though replication across seeds remains uncertain.

  • Limitations: Replication across seeds is uncertain because this prompt shares the structural family of the seed-2 collapse, while the comparison also uses different decoding temperatures.The frontier baseline uses temperature 0.3, whereas the trained model decodes greedily.
  • Original-case transfer: 0% commitment at every level on the original 40 cases demonstrates transfer of the trained gate.Table 5 reports commitment on the original benchmark, using the same cases, prompt, and commitment definition as the frontier comparison.
  • Original-case transfer: At L1, tool calls fall from 27 of 40 for the fine-tuned model to 1 of 40 for the preference-trained variant.The remaining L1 tool calls are identified as the wrong kind of refusal, and Cohen’s h at L2 is +1.65.
  • Original-case transfer: 100% of matched answerable questions are answered with no wrong directional call, showing that the decline is selective.The apparent scoring errors are exactly-50% hedges; excluding those hedges, directional accuracy is 100% at L1 on 36 of 40 items.

8 Where the trained gate breaks

The trained gate holds when prompts allow reasoning but breaks under rigid, reasoning-suppressing formats, where performance and calibration deteriorate. Its apparent failures also vary across runs, indicating that deployment requires testing response-format boundaries and training stability rather than relying on a single aggregate metric.

  • Format boundary: Under the main 540-case recipe, commitment on unknowable items remains low under the reasoning-suppressing framing, reaching 16.7% and 8.3% on crypto and never exceeding 4.2% on sports or weather.When Youden’s J declines, the model often substitutes a tool call that cannot observe a future event, so refusal remains high despite poor discrimination.
  • Run-to-run instability: The 516-case ablation recipe is unstable under the same framing: three of four runs commit at L2, ranging from 12 committed items in seed 2 to none in seed 3, with seeds 0 and 1 the heaviest cases.Reporting both commitment and Youden’s J is necessary because a zero J can reflect universal refusal or universal commitment; an earlier GRPO attempt showed the same run-to-run instability, suggesting a capacity limit at this model scale.
  • Format boundary: Reasoning-permitting prompts preserve strong gate discrimination, whereas a structurally different tool-offering format ranges from near-perfect to no discrimination across runs.The original and natural framings provide reasoning slots; the frontier transfer framing requests only a decision and probability. Figure 6 shows that Youden’s J alone can conceal whether low discrimination reflects refusal or commitment.
  • Rigid output formats: Rigid <answer>-tag formats eliminate reasoning and reduce answerable-question accuracy to 50.9-73.8%, versus 88.9-98.6% under natural framing on the same questions.The rigid format produces a bare two-line response, while natural framing produces a reasoning block in every response; the supplied passages caution that format and reasoning are in tension.
  • Calibration failure: Under the rigid format, mean stated confidence is 84.0% despite 57.2% accuracy on answerable questions, while added evidence shifts YES probabilities toward hedging from {50: 4, 60: 2, 80: 1, 90: 4} to {50: 9, 60: 1, 70: 1}.The trained model’s stated number therefore no longer tracks correctness reliably: rigid formatting relocates miscalibration, and more evidence increases hedging on answerable questions.

9 Discussion

The discussion locates the failure in an action gate bypassed by authoritative presentation, not in belief calibration or judgment. It argues that commitment on irreducible probes provides a practical governance metric while preserving answerability on matched knowable questions.

  • Action calibration: A 48-point action shift accompanies only a 3-point probability shift, showing that stated probabilities do not calibrate action quality.Probabilities also rank correctly answered events below incorrectly answered ones, while nominally justified actions score worse than silence and training inflates probability in the correct-action cell.
  • Presentation effects: Authoritative presentation, rather than information volume, licenses action and erodes agents’ willingness to acknowledge irreducible uncertainty.The scrambled control indicates that the mechanism is not simply overweighting noisy evidence.
  • Failure mechanism: The failure is a bypassed act-or-decline gate, while judgment remains intact and commitment produces herded, momentum-shaped calls worth less than silence.Its localization makes the failure addressable by instruction or training: 540 synthetic dice-and-coin examples moved commitment from 54% to 0%.
  • Governance measurement: Commitment rate on aleatoric probes is a cheap, reproducible governance test that avoids rewarding agents merely for answering more questions.Auditing only answerable questions rates the L2 agent higher despite worse decision quality than silence; the matched answerable arm makes the metric non-gameable.

10 Related work

Related work distinguishes abstention from forecasting quality and separates knowing from acting. This paper contributes a causal behavioral localization of evidence-induced commitment on aleatorically unknowable future events, where the evidence is constructed to be empty.

  • Aleatoric versus epistemic unanswerability: AbstentionBench (Kirichenko et al., 2025), Abstain-R1 (Zhai et al., 2026), and TruthRL (Wei et al., 2026) study missing-information or ill-posed questions, whereas this work targets future-event unpredictability.AbstentionBench reports that reasoning fine-tuning degrades abstention; Abstain-R1 adds post-refusal clarification for queries clear in meaning but unresolved from available information.
  • Forecasting evaluation: ForecastBench (Karger et al., 2024) evaluates forecast quality against resolved outcomes and reports systematic overconfidence, whereas this work measures whether an agent forecasts at all.The study borrows forecasting benchmarks’ temporal-separation control: every as-of date follows every roster model’s training cutoff, preventing leakage.
  • Knowable versus unknowable: Ahdritz et al. (2024) probe aleatoric language ambiguity internally, while this work provides a black-box behavioral complement on resolved future events.Here, aleatoric unanswerability concerns unpredictable real-world outcomes rather than next-token ambiguity.
  • Knowing versus acting: Sun et al. (2026) decode tool necessity at AUROC 0.89-0.96, while this work behaviorally localizes a causal knowing-versus-acting gap through evidence presentation and sealed-outcome evaluation.Kadavath et al. (2022) showed models are largely calibrated about their own knowledge; Sun et al. (2026) found hidden-state tool-necessity signals substantially exceeding verbalized reasoning.
  • Evidence-induced miscalibration: Related accounts connect the pattern to evidence-induced confidence, guessing incentives, distractor effects, and human illusion of validity, but this work uses empty evidence and measures action.Xuan et al. (2026) attribute inflated verbalized confidence to retrieval noise; Kalai et al. (2025) argue evaluation rewards guessing; human analogues include Oskamp (1965), Slovic & Corrigan (1973), and Tversky & Kahneman (1973).

11 Limitations

The study closes two experimental threats but leaves six limitations concerning domain validity, model heterogeneity, lexical confounds, intervention scope, measurement quality, and reproducibility. These constraints motivate future tests across agent loops, model scales, domains, and developer model families.

  • Weather is a weak instrument because its real ten-day ensemble rain probabilities have genuine skill and lack sealed outcomes, making correct answers difficult to distinguish from seduction.Crypto, which has both sealed outcomes and the relevant properties, is the primary transfer domain.
  • Three of twelve models show the effect, while four never commit and three always do, so the pooled rate does not generalize to frontier models broadly.The analysis rests on models from one developer.
  • A residual lexical confound remains because synthetic unknowable items explicitly name randomness, although transfer panels omit that vocabulary and the gate still fires.The tense confound and training-set overlap with one transfer domain were closed experimentally, but six threats remain open.
  • The intervention is demonstrated only on one 3B architecture with synthetic data, so its survival at larger scales remains unestablished.The study used six independent runs of the main recipe and four of the ablation, with 24 to 40 cases per cell and one sample per cell.
  • 52.8% of Qwen3.7-plus’s L2 unknowable responses and 23.6% of Llama 3.3 70B’s are unparseable, making their discrimination scores lower bounds and qualifying the reported per-model spread.Unparsed responses count as non-commitments; the corresponding J values are +17 and +53, compared with 0.0–2.8% unparseable responses for the other ten models.
  • Individual cell values can shift by one to two cases out of 24 because greedy decoding is not bit-reproducible, and hosted-model serving configurations are uncontrolled.The reported values should therefore be treated as approximate.
  • Future work should test commitment propagation through agent loops, a second model scale, domains where declining is unsafe, and differences among models within one developer family.

12 Conclusion … Appendix C - Parser pair and the strict/semantic delta

The study locates the failure in a trainable but prompt-fragile act/don’t-act gate: fabricated evidence can trigger action, while training restores restraint without removing answerability. The appendices document preregistration boundaries, prompt variants, and parser checks that affect interpretation.

  • 12 Conclusion: On unanswerable questions, agents refused bare prompts but committed with professional-looking panels, including fabricated ones; probabilities were anti-predictive (AUROC 0.346) and calls worth less than silence.The judgment needed to refuse was present and elicitable, despite the action failure.
  • 12 Conclusion: A 3B model trained on dice, coins, and timers stopped committing on stock and crypto forecasts while preserving answerable responses, but prompt-shape changes made the gate fragile.The conclusion characterizes this as a gate installable by 540 synthetic examples, not a missing capability.
  • Appendix A - Pre-registration and analysis decisions: The preregistration covered the evidence gradient, L2’ control, mitigation, truthfulness metric T, +42 baseline, and matched answerable controls before the relevant analyses.The answerable controls were intended to distinguish a knowability gate from blanket refusal.
  • Appendix A - Pre-registration and analysis decisions: The authors label the coverage ablation, tense-balanced control, semantic parser, and ±5-point TOST margin according to when each decision was made, including post-hoc choices.They also report predictions that failed, including expectations about preference training and removing an overlapping training family.
  • Appendix B - Prompts (verbatim): The evidence-gradient prompts kept the response format constant, while thin, rich, and scrambled arms varied the displayed indicators, with scrambled fields taken from an earlier date.Training used ten framings with or without a tool; DECLINE or ANSWER remained gold because no search resolves a random future.
  • Appendix B - Prompts (verbatim): Held-out evaluation used a natural framing with new wording and a frontier framing that retained the original three-option agentic structure with a tool call.Both framings requested a parseable decision line.
  • Appendix C - Parser pair and the strict/semantic delta: The semantic parser accepted equivalent decision prefixes beyond the strict DECISION: line; parser results matched everywhere except one preference-trained checkpoint, preventing format compliance from hiding a working model.The strict parser accepts only DECISION:, whereas the semantic parser also accepts RESPONSE:, FINAL DECISION:, VERDICT:, ACTION:, CHOICE:, and ANS:.

Appendix D - Per-model results

At the heaviest evidence level, models answer answerable items essentially perfectly, so per-model variation in J is driven by whether they decline unknowable items. Two models’ J values are substantially affected by unreadable responses that parsers count as non-commitments.

  • Per-model results: Near-100% answerable accuracy and zero declines on answerable items leave unknowable-item decline rates as the source of variation in J.This pattern holds across crypto, sports, and weather transfer domains, with 72 unknowable and 72 answerable decisions per model.
  • Per-model results: 52.8% of Qwen3.7-plus’s and 23.6% of Llama 3.3 70B’s L2 unknowable responses are unparsed, versus 0.0–2.8% for every other model.Because unparsed responses count as non-commitments, the two models’ J values (+17 and +53) are computed on a minority or ba

Appendix E - Training details · Appendix F - Transcripts

Training used balanced synthetic cases with deliberately varied prompt framings and held-out domains, while transcripts show both fabricated-panel commitment and correct refusal under explicit uncertainty. The appendices also document format-induced suppression of reasoning and scoring limitations affecting interpretation.

  • Appendix E - Training details: 540 balanced synthetic cases used verified ground truth, matched predictive and non-predictive panels, and six families held out entirely.The ablation used 516 cases.
  • Appendix E - Training details: Prompt framings were stratified to emphasize tool-offering seductive panels, while half were enriched into multi-line professional formats matching transfer-panel density.The tool-offering seductive-panel cell comprised 25.5% of training rows versus 9.0% under random sampling.
  • Appendix E - Training details: Training used 4-bit QLoRA with rank 32, alpha 64, three epochs, learning rate 2e-4, completion-only loss, and greedy 256-token evaluation.The setup used nf4 double quantization, attention projections, batch 8, and gradient accumulation 2.
  • Appendix E - Training details: A fifth independent run was evaluated but excluded from transfer-domain transcripts because its generations were not retained; its tense-control generations remained available.Those retained tense-control generations are reported in §6.6.
  • Appendix F - Transcripts: Transcripts show a model committing on a fabricated Bitcoin panel whose cited indicators were real numbers from another date.The example presents a professional-looking technical-panel rationale despite fabricated evidence alignment.
  • Appendix F - Transcripts: The same models correctly refuse an unpredictable trading question and answer near-perfectly when identical irreducible uncertainty is explicitly labeled as a fair coin.The transcript states that a fair coin has no memory, so prior flips and commentary do not matter.
  • Appendix F - Transcripts: The inherited scoring rule penalized a numerically neutral 50% hedge even when its reasoning was correct and explicit on a YES item.The transcript’s probability output expressed no view, but was still counted as an error.
  • Appendix F - Transcripts: Rigid wrapper tags suppressed reasoning entirely in 288 of 288 main-run responses, whereas natural framing elicited reasoning in all 288 responses.Across eleven checkpoints with retained generations, the wrapper-tag condition suppressed reasoning in 3164 of 3168 responses.
Loading 2608.27167v1…