Source-linked AI summary
Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking
Zhexi Feng, Wuxi Chen, Bingrui Zhang
TL;DR
Finite-reference evaluation can mark valid open-ended ToM beliefs false, raising a question about whether confidence-rule rankings remain trustworthy. The paper fixes emitted beliefs and paired scores, compares reference-derived with human literal-truth labels, and develops a pilot-anchored restoration method. It finds repeated ranking reversals, including in a released NQ-open pipeline, and uses TriSource-Restore to repair confidence subject to deployment safeguards.
Problem
Finite references and matchers can label valid open-ended beliefs as false, creating a gap between reference-derived evaluation labels and literal-truth labels.
Method
The paper holds emitted contents and paired confidence rules fixed, decomposes label-source distortion, and anchors auxiliary judgments to a probability-sampled human pilot through TriSource-Restore.
Results
90–96% of reference-unmatched OpenToM beliefs were literally true, while strictly proper Brier rankings reversed across all six authored scenarios and ICE rankings reversed on 301 NQ-open predictions.
Takeaways & Limitations
Reference-derived labels can reverse model selection and calibration conclusions even when outputs and scores are unchanged, motivating human-anchored auditing and repair.
Takeaways & Limitations
Literal truth is the target, while usefulness, coverage, and false commitment require separate measures; the controlled case-study scripts and gold are project-authored stress conditions, not a random sample.
Abstract
from arXiv · showhide
Open-ended Theory-of-Mind (ToM) trackers emit valid beliefs absent from finite references. A finite-reference-plus-matcher pipeline marks unmatched outputs false, creating proxy labels that can reverse proper-score model selection on fixed outputs. Holding 259 beliefs and paired scores fixed, reference recoding lowers weighted prevalence from 0.783 to 0.295 and reverses strictly proper Brier risk: a frozen source-prior rule leads native confidence by 0.227 under reference labels and trails by 0.152 under blinded adjudication, in all six authored scenarios. A reference-only Platt recalibrator reverses further. An ICE-specific reversal appears in a released 301-question NQ-open DPR-BERT pipeline: its average-confidence baseline improves instance-level calibration error by 0.045 under exact match but worsens it by 0.074 under human correctness, with both intervals excluding zero. On independently authored OpenToM narratives, 90-96% of audited unmatched beliefs are literally true and the paired direction again reverses. An exact decomposition attributes the distortion to omitted truths, and a closed-form criterion correctly classifies comparisons from twelve released systems. Frozen-audit retrospective replay shows 50 attempted annotations recover ranking direction with probability at least 0.996. TriSource-Restore anchors full-frame reference labels and frozen automatic judgments to a probability-sampled human pilot, maintains at least nominal coverage, narrows intervals, and repairs confidence subject to a base-rate deployment gate.
1 Introduction
Open-ended ToM trackers generate propositions beyond finite reference sets, so unmatched outputs may be valid rather than false. The paper shows that reference-derived labels can reverse confidence-rule rankings and proposes auditing and repair.
- Open-ended ToM trackers maintain fine-grained propositions as evidence arrives, unlike completed-narrative tests with fixed question sets.
- Finite references can omit valid micro-beliefs, while unmatched-as-false labels penalize justified confidence and can rank the same confidence rules oppositely.
- 0.227 lower Brier risk under finite-reference labels became 0.152 higher risk under adjudicated literal-truth labels for the frozen source-prior rule, across all six scenarios.The comparison holds emitted contents and confidence rules fixed while changing only the label source.
- 90–96% of reference-unmatched beliefs were literally true in the frozen OpenToM audit, with the paired ranking direction again opposed.The audit covered 240 attempted items, of which 209 were usable.
- 301 frozen NQ-open predictions from a released DPR–BERT pipeline exhibited a significant ICE ranking reversal between exact-match and human-correctness labels.
- TriSource-Restore anchors full-frame reference labels and frozen automatic judgments to a probability-sampled human pilot, narrowing intervals and escalating when uncertainty remains.Fifty attempted annotations safely recover ranking direction in all three audited units.
2 Related Work
Prior work addresses fixed task items, evidence construction, missing labels, and human-defined estimands, but does not study persistent tracker propositions scored false solely because they are absent from finite references. This paper connects that gap to released open-domain QA systems and a pilot-based restoration procedure.
- ToM benchmarks generally score beliefs in completed narratives or fixed question sets, whereas this paper studies reference-derived labels for tracker-generated beliefs about unreliable agents.
- Open-ended content verification and information-retrieval evaluation address evidence construction or unjudged items, but do not study persistent tracker propositions marked false solely for finite-reference absence.
- Released open-domain QA systems combine free-text predictions with finite-gold exact-match labels, creating the setting in which calibration rankings can invert.
- Existing calibration and missing-label methods typically presuppose a defined example set and observed target, while prediction-powered inference preserves a human-defined estimand using smaller labeled samples.
- The paper’s operating procedure targets a paired proper-score gap with transfer-safe auxiliary weights, a residual gate, and a pilot-driven repair and escalation rule.
3 Problem and Method
The method defines literal truth as the confidence target, isolates label-source effects on identical beliefs, decomposes proper-score distortion, and uses pilot-anchored auxiliaries for restoration.
- Problem and target: Confidence c is operationalized as a probabilistic score for whether a fixed emitted proposition is literally true in the scenario, denoted Y_b = 1.Reference overlap, task relevance, and downstream usefulness are treated as different estimands.
- Label-source decomposition: Z combines finite-reference coverage with automatic matching, so Y = 1 and Z = 0 may reflect omitted references, matcher false negatives, or granularity mismatch.
- Label-source decomposition: A proper score decomposes ranking distortion through omitted truths and false-positive matches; the closed-form reversal condition is structural for Brier risk but not for nondecomposable ECE.
- Identification: The identification holds propositions, targets, provenance, revision histories, sampling weights, and belief contents fixed while changing confidence rules or label source.
- Restoration: TriSource-Restore estimates the human-truth paired gap using full-frame labels, frozen confidence-blind automatic probabilities, and probability-sampled human labels.Auxiliary weights are frozen independently of the target pilot, with nonnegative weights summing to at most one and a residual-MSE improvement gate.
- Restoration: A monotone Platt map repairs confidence by combining auxiliary log loss with inverse-probability-weighted pilot correction to human truth.
- Controlled case study: The controlled case study uses tracker propositions with native confidence, provenance, and revision history, while EG replaces only confidence with a frozen source prior before applying shared revision multipliers.
4 Benchmark and Evaluation
The benchmark separates coverage, truth calibration, discrimination, and commitment while evaluating beliefs against finite references and blinded human judgments. Its audits use frozen samples, independent volunteers, clustered inference, and distinct annotation budgets.
- Benchmark population: DoubtfulToM-Bench contains six hand-authored interactions with sequential events and 371 task-defined true and false propositions.Its gold labels encode scenario facts but cannot enumerate every reasonable micro-inference or attributed observation.
- Evaluation protocol: Evaluation keeps coverage, truth calibration, discrimination, and commitment as separate metrics and uses confidence-blind matching.Reference-relative ECE diagnoses protocols rather than truth calibration.
- Evaluation protocol: Brier is the primary paired statistic because it is strictly proper and unbinned, while ECE and AUROC describe aggregate calibration and discrimination.Reference omission can lower the weighted positive-label rate and change ECE through its first moment.
- Human audits: Three external volunteers independently judge sampled beliefs as true, false, or insufficient using complete context rather than reference resemblance.They are blind to method, confidence, source, matching, and hypotheses where applicable.
- Human audits: 259 of 280 local annotated items and 209 of 240 OpenToM items yield canonical usable labels.The external OpenToM sample and paired intervention were frozen before annotation.
- Pilot validation: Operational replay uses 50 attempted annotations, whereas TriSource-Restore fixes 50 usable truth labels to evaluate RMSE, interval width, coverage, and stopping.Downstream repair holds out test clusters from selection, fitting, and transfer weights.
5 Results
Across controlled audits and released benchmarks, finite-reference labels reverse calibration and proper-score rankings when valid unmatched beliefs are treated as false. The reversal is driven mainly by omitted truths and prevalence collapse, persists across datasets and scoring protocols, and is predictable from a closed-form criterion.
- 5.1 A Released NQ-open Pipeline Reverses Its Ranking: 115/301 predictions were exact-match correct versus 160/301 under human judgment, reversing the released pipeline’s average-confidence ICE ranking from −.045 to +.074.Both intervals excluded zero; Brier gaps reversed similarly but the prespecified claim concerns ICE.
- 5.1 A Released NQ-open Pipeline Reverses Its Ranking: 9.9 pp higher human accuracy reversed the LCCmain2002-versus-InstructGPT ordering in an independent 444-question CuratedTREC audit.The audit used finite regex patterns and human judgments.
- 5.1 A Released NQ-open Pipeline Reverses Its Ranking: AUROC .663 and DSC .0302 measured cross-system agreement across 1,490 human-labelled answers, while the crossover criterion classified all 13 reversals and 8 non-reversals.A reference-only probe reached DSC .107 against adjudicated labels.
- 5.2 Adjudicated Literal Truth Reverses Proper-Score Risk: .227 lower Brier risk under finite-reference labels became .152 higher risk under adjudicated labels for the frozen source-prior probe on 259 beliefs.The direction reversed in all six authored scenarios; reference prevalence fell from .783 to .295.
- 5.2 Adjudicated Literal Truth Reverses Proper-Score Risk: Reference-only Platt recalibration worsened adjudicated Brier risk from .178 to .408 while improving reference-label risk from .446 to .205, reversing all six scenarios.It transferred the reference prevalence collapse into deployed scores without seeing adjudicated labels.
- 5.2 Adjudicated Literal Truth Reverses Proper-Score Risk: Closed-world scoring agreed with adjudication on 1,855 propositions, whereas unmatched-as-false alone selected the probe under both Brier and ECE.Matched-only scoring selected native confidence but could not estimate absolute calibration.
- 5.3 Omitted Truths Drive the Distortion: Omitted truths dominated all three audited units, with an audit-conditioned restoration threshold q* ranging from .570 to .654.Fixed-seed Bernoulli restoration matched the analytic curves.
- 5.3 Omitted Truths Drive the Distortion: 73% of reference-unmatched local beliefs were true, and 70% of unmatched-true weight had native confidence ≥.8.The authors attribute the headline to the full reference-derived label pipeline because matcher false negatives were not separately estimated.
A. Ranking inference (50 usable truth labels; 5,000 repetitions)
The section reports 95% interval width and H/T coverage, with separate values for Local, OpenToM-V, and OpenToM-C.
- 95% width and H/T coverage are the reported evaluation quantities.
- Local is reported with paired values .0483/.0309, .2146/.1356, .959, and 36.8%.
- OpenToM-V is reported with paired values .0134/.0113, .0888/.0833, .996, and 6.1%.
- OpenToM-C is reported with paired values .0212/.0188, .1510/.1009, .961, and 33.2%.
B. Population-weighted held-out Brier risk (lower is better)
The held-out Brier-risk comparison reports TriSource-Restore against human-only and full-training-Y references for Local, OpenToM-V, and OpenToM-C settings.
- The reported rows compare Local, OpenToM-V, and OpenToM-C across paired values for the listed methods.
- Held-out Brier risk is reported in Panel B, with lower values preferred.V/C denote OpenToM Vanilla/Conservative, while full-training-Y is an empirical comparator rather than an optimization bound.
- Reference-relative checks do not establish whether labels track truth because they condition on the same finite reference.The paper separates calibration and coverage and makes no system-coverage claim.
6 Discussion and Audit Implications
The discussion frames the findings as a protocol-level ranking failure attributable to label source, while separating ranking identification, deployment gates, and coverage claims.
- Fixed scored objects across NQ-open, authored scenarios, and OpenToM support a protocol-level failure rather than universal tracker or matcher coverage.
- The estimand is literal correctness of emitted beliefs, and paired confidence rules on fixed contents isolate the label source as the changing factor.
- If the interval crosses zero or fewer than eight clusters are sampled, the protocol requires more human labels rather than termination.
- TriSource-Restore is an offline evaluation-time repair: its auxiliary Q is a control variate, not a human substitute.
7 Conclusion
The conclusion states that unmatched beliefs need not be false and that finite-reference labels can reverse rankings on fixed emitted contents.
- Finite-reference matching can mark valid unmatched beliefs false and reverse strictly proper Brier rankings on identical emitted contents.
- The released NQ-open pipeline exhibits the same label-source failure.
- A 50-attempt human pilot recovers the ranking direction on all three audited units.