Source-linked AI summary
Ordinal Gates, Cardinal Bets: Matching LLM Confidence to the Financial Decision Operator
Rayansh Singh, Sara Rezaeimanesh
TL;DR
The paper asks whether LLM confidence scores can be evaluated independently of the decision operators and exposure controllers that consume them. Using FactSet news and nine open-weight LLMs, it tests matched and mismatched confidence maps and scales, finding that matching improves frozen-scale performance but the gain attenuates under adaptive risk control and later evaluation.
Problem
LLM confidence scores have operator-relative decision value: ordinal gates consume ranking, whereas cardinal sizing consumes magnitude and depends on its paired controller.
Method
The study fits confidence maps and exposure scales on 2021 FactSet news for Nasdaq-100 equities and evaluates nine open-weight LLMs out of sample on 2022–2023, including ordinal gates and cardinal sizing.
Results
Matching each confidence map with its fitted scale improves ensemble CER by 9.2pp/yr under frozen-scale control, while the benefit falls to 1.6pp/yr under identical adaptive-volatility control.
Takeaways & Limitations
Confidence transformations should be monitored and revalidated jointly with the downstream portfolio-construction controllers that consume them.
Takeaways & Limitations
The evaluation has limited generalizability: only 3 of 9 per-model effects individually survive Holm, and primary-period gains are concentrated in 2022 and non-technology tickers.
Abstract
from arXiv · showhide
LLM confidence scores are not independently deployable objects: their decision value depends on the downstream operator and exposure controller that consume them. Monotone recalibration cannot change a coverage-matched rank-based gate, whereas position sizing consumes score magnitude, so changing a confidence map can invalidate a scale fitted to the previous score distribution. We test this on FactSet news for Nasdaq-100 equities, fitting maps and scales on 2021 and evaluating nine open-weight LLMs out-of-sample on 2022--2023. Cross-applying raw and correctness maps with independently fitted scales shows that the two components are not portable alone: scale transfer reduces certainty-equivalent return (CER) in $8/9$ models and produces large risk-target errors. Matching each map with its fitted scale improves ensemble CER by $9.2$ percentage points per year under frozen-scale control ($p<0.001$), and the effect remains significant when the single largest-contributing model is excluded ($+5.5$pp/yr), so it is not driven by one case. Under an identical adaptive-volatility controller, however, the incremental effect falls to $+1.6$pp/yr, with a significant controller interaction. Annual walk-forward effects are smaller, although map--scale interaction remains positive in every fold. Confidence transformations should therefore be evaluated jointly with the downstream controllers that consume them.
1 Introduction
The paper argues that LLM confidence has decision value only relative to the downstream operator and exposure controller. It formalizes the ordinal–cardinal distinction and tests whether confidence maps and fitted scales must be paired.
- Prior work studied directional financial signals, confidence miscalibration, reliability, and threshold selection without isolating which score property each decision requires.
- Ordinal operators use confidence rankings, whereas cardinal operators use magnitude and therefore depend on the paired controller.The distinction covers thresholding, selection, and routing versus sizing, allocation, and portfolio weighting.
- The study evaluates FactSet Nasdaq-100 news from 2021–2023, fitting maps and controls on 2021 and testing forward on 2022–2023 across nine open-weight LLMs.The ordinal analysis also uses five confidence-extraction channels across two open-weight models.
- 9.2pp/yr: pairing each confidence map with its fitted scale improves ensemble CER under frozen-scale control, while the gain largely dissolves under identical adaptive-volatility control.Cross-applying maps and independently fitted scales produces large risk distortions.
- Confidence maps and fitted exposure controls are coupled rather than independently portable, making their joint evaluation the paper’s general design implication.
2 Related Work
Related work distinguishes statistical calibration and selective decisions from this paper’s narrower focus on operator-relative confidence. The paper connects monotone calibration’s rank preservation to financial portfolio decisions.
- Calibration research commonly evaluates reliability with ECE or Brier score and post-hoc scaling, while selective decisions consume rankings or thresholds.
- Monotone calibration cannot reorder predictions, so it cannot change a coverage-matched selective-classification gate even if it improves reliability.
- The paper’s narrower claim is that financial operators consume different properties of one signal: rank for ordinal decisions and magnitude for cardinal exposure control.
- FinDPO exemplifies the gap by temperature-scaling saturated sentiment scores before applying an ordinal fixed-quantile portfolio rule.
3 Theory
The theory separates rank-invariant gates from magnitude-sensitive sizing. It defines map–controller compatibility and explains why a changed confidence shape can require a refitted exposure scale.
- 3.2 Cardinal operators: map-specific scales and magnitude sensitivity: The paper defines a chain from confidence map to shape, scalar scale, controller, and system, then tests whether maps and fitted scales are compatible.
- 3.1 Ordinal operators: a resolution bound on gate accuracy: A coverage gate retains the top-φ fraction by score, and its accuracy gain is tied to confidence–correctness covariance.
- 3.1 Ordinal operators: a resolution bound on gate accuracy: Strictly increasing transformations preserve threshold families, coverage-matched portfolios, and Brier resolution, whereas fixed numerical thresholds remain magnitude-sensitive.
- 3.2 Cardinal operators: map-specific scales and magnitude sensitivity: Cardinal sizing should target conditional signed payoff relative to risk rather than correctness probability alone; correctness is used as a lower-variance proxy in the implemented system.
- 3.2 Cardinal operators: map-specific scales and magnitude sensitivity: Map-specific scales arise because the implemented sizing rule multiplies a map-specific shape by one scalar, so transferring another shape’s optimum incurs regret.
- 3.2 Cardinal operators: map-specific scales and magnitude sensitivity: Uniform positive rescaling cannot explain the matched effect because refitting offsets it before caps bind; nonlinear reshaping, support changes, cap interactions, or shift are required.
4 Data and Evaluation Protocol
The evaluation uses frozen, look-ahead-free calibration and testing on FactSet news for Nasdaq-100 equities, with separate ordinal and cardinal decision protocols. Portfolio construction applies fitted maps, exposure scaling, caps, costs, and pre-specified inference to assess economic outcomes.
- Data and splits: 2021 calibration covers 15,034 ticker-days, while 2022–2023 evaluation covers 31,888 ticker-days across 61 frozen Nasdaq-100 tickers.Headlines are aggregated by ticker-day and assigned to the first trading session after publication.
- Labels and horizons: Five-session cardinal decisions use a holding-period correctness target, whereas ordinal decisions evaluate the predicted direction against the next session’s open-to-close return sign.The five-session horizon was fixed a priori, with sensitivity effects remaining positive through h=8.
- Models and confidence signals: The cardinal grid contains nine open-weight checkpoints spanning 7–32B parameters and five model families, using verbalized confidence across models.The evaluated families are Qwen, Gemma, Mistral, Phi-4, and FinLLaMa.
- Portfolio construction: For each map, isotonic calibration produces confidence-to-success estimates that feed signed positions, capped aggregate holdings, and portfolio returns across five overlapping entry cohorts.Positions are clipped to a 10% per-name cap and rescaled to a 100% gross cap when necessary.
- Inference and estimands: Exposure scales are selected on 2021 returns to target 5% annualized pre-cost volatility, while all maps, thresholds, and scales remain frozen during evaluation.CER with gamma=3 is the primary economic outcome, tested with a joint circular block bootstrap and multiple-testing correction.
- Observed portfolio behavior: The matched system reduces condition-averaged per-model volatility from 9.9% to 8.3% and ensemble volatility from 8.8% to 7.8%.Average daily active positions also fall from 18.1 to 14.5, while portfolio constraints are not binding.
- Inference and estimands: The primary economic estimand is the matched-versus-raw ensemble effect under frozen-scale control, with controller choice, costs, exclusions, and annual transfer treated as robustness checks.The estimand concerns the fixed ensemble on the observed path rather than a population of LLMs or market regimes.
5 Ordinal Decisions
Ordinal recalibration can improve statistical reliability without changing coverage-matched decisions. However, the tested confidence signals have weak resolution, so neither selective prediction nor fixed-threshold gating produces reliable out-of-sample gains.
- Calibration without decision change: Logit ECE falls from approximately 0.47 to 0.03–0.04 after temperature scaling, while the induced ranking and coverage-matched gate remain unchanged.Softmax and temperature-scaled sigmoid are both strictly increasing in the logit difference.
- Resolution and selective gains: Resolution is only 0.07%–0.4% of uncertainty across 18 configurations, limiting the accuracy improvement available to coverage gates.At the 10% coverage floor, the bound permits roughly 4–9 percentage points even under best-resolution conditions.
- Resolution and selective gains: The base classifier reaches only 50.3–51.8% one-session accuracy, at or below the 53.8% base rate, and none of five confidence channels provides dependable selection.A K=5 self-consistency baseline is unanimous on 96–98% of ticker-days and adds no resolution.
- Fixed-threshold gating: Zero of 18 accuracy-tuned and zero of 18 CER-tuned fixed-threshold configurations survive family-wise multiple-testing correction.Although 15 of 18 CER-tuned configurations have positive point estimates, the improvements are not statistically reliable.
6 Cardinal Decisions
Cardinal sizing consumes confidence magnitude, so maps and exposure scales are not independently portable. Matching each correctness map with its fitted scale improves frozen-scale CER, but the gain is smaller under adaptive volatility and annual transfer.
- Portability diagnosis: 8/9 models: applying the correctness-map scale to the raw map reduces CER, showing that scales are not portable alone.Transferred scales also produce substantial realized-volatility errors relative to their intended map.
- Matched pairing: 9.2pp/yr: matched map–scale pairing improves ensemble CER under frozen-scale control.The joint effect is significant at p<0.001 and remains +5.5pp/yr after excluding Qwen2.5-7B.
- Matched pairing: +9.0pp/yr of the +9.2pp/yr matched effect comes from mean returns, while +0.2pp/yr comes from the variance component.The matched improvement is positive in all nine models, and realized volatility falls in all nine.
- Controller boundary: +8.2 versus +6.6pp/yr: the correctness map still exceeds the raw map under the identical adaptive-volatility controller.The incremental advantage is limited and statistically imprecise, with a significant interaction favoring frozen-scale control.
- Mechanism and robustness: Most of the matched gain survives common trade support, indicating magnitude reallocation rather than filtering drives the result.The matched effect remains positive across alternative raw sizing maps and seven model-family compositions.
- Temporal transfer: Annual walk-forward gains are attenuated because positive map–scale interaction is offset by adverse scale transfer, although the interaction remains positive in every fold.Under isotonic calibration, the pooled components are −4.0pp/yr from map replacement, −8.1pp/yr from scale transfer, and +13.3pp/yr from interaction.
7 Limitations and Conclusion
The paper limits its conclusions by weak and potentially contaminated signals, partial generalizability, and temporal instability, while concluding that confidence maps and exposure scales should be evaluated as a coupled portfolio-construction system.
- Limitations: The base classifier achieves 50.3–51.8% accuracy against a 53.8% base rate, so the paper makes no market-beating claim.Pooled diagnostics show fitted correctness and signed payoff generally rise together, but not within every model or masking condition.
- Limitations: All nine models were released after the primary 2022–2023 evaluation, leaving contamination from pretraining exposure incompletely ruled out.The authors state contamination could inflate the headline effect, while the map–scale coupling result does not depend on any period’s return sign or magnitude.
- Limitations: Only 3 of 9 per-model effects survive Holm correction, and primary-period gains concentrate in 2022 and non-technology tickers.The reported subgroup effects are +16.2pp/yr for 2022 and +13.9pp/yr for non-technology names, so the result is partial rather than universal.
- Limitations: Annual walk-forward gains are smaller and do not survive family-wise correction, although the map–scale interaction remains positive across folds.The temporal extension therefore provides weaker evidence for effect magnitude than the primary evaluation.
- Conclusion: Monotone recalibration preserves coverage-matched gates, but changing a confidence map can alter portfolio exposures and invalidate a scale fitted to the original distribution.The conclusion recommends monitoring and revalidating confidence maps as part of the full portfolio-construction system.