Source-linked AI summary
Dutch Books for Language Models
Isaiah Andrews, Suproteem Sarkar
TL;DR
The paper asks whether language-model probabilistic forecasts are internally coherent, an important issue when users rely on them for uncertain decisions. It tests coherence with de Finetti-based Dutch-book linear programs over stock-return events, finding substantial incoherence that worsens with richer logical relationships and can rise sharply with irrelevant context.
Problem
Language models are increasingly used for probabilistic forecasting, but evidence shows they can violate consistency properties, motivating a direct coherence evaluation.
Method
The paper elicits forecasts over logically related stock-return events and computes the largest guaranteed Dutch-book profit using a linear-program measure of incoherence.
Results
Incoherence is common and often sizable, increases with richer logical relationships, and can rise by an order of magnitude when irrelevant contextual details are added.
Takeaways & Limitations
Coherence can be evaluated without outcome labels, and the results motivate training strategies that encourage computation across related events.
Takeaways & Limitations
The evidence comes from variance-normalized stock-return forecasting and elicitation experiments based on GPT-OSS-120B, so generalizability and invariant cross-model ordering are not established.
Abstract
from arXiv · showhide
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that these forecasts fall out of a coherent world model. In this paper, we evaluate the coherence of language model probabilistic forecasts through a procedure that builds on a theorem due to de Finetti. We elicit forecasts from language models across events generated from stock returns data. We then use linear programs to compute the largest Dutch-book profit - the profit an arbitrageur could guarantee by betting against model-generated probabilities - which we use as a measure of incoherence. Our procedure does not require outcome labels, so we can evaluate coherence even in settings where outcomes are not observed or have not yet resolved. We find substantial evidence of incoherence in language model forecasts. Such incoherence increases when there are richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude. We conclude by discussing how alternative training strategies may improve probabilistic coherence.
1 Introduction
The paper develops a global, arbitrage-based test of whether language-model probability forecasts are coherent, evaluating it on logically related stock-return events. It finds incoherence is common, increases with richer event relationships, and can be strongly affected by elicitation details.
- Language models increasingly provide probabilistic forecasts, but may violate basic consistency constraints such as event-complement probabilities summing to one.
- The method uses de Finetti’s theorem to identify incoherence through the largest profit an arbitrageur can guarantee against stated probabilities.
- The evaluation constructs related events from stock-return bins, unions, and complements across days and assets.
- The procedure is exhaustive for a fixed event set and does not require outcome labels.
- Incoherence is common and often sizable, with larger variation across forecasts than accuracy shows.
- Richer logical relationships increase incoherence, while irrelevant contextual details can increase it by an order of magnitude.
- The paper discusses training strategies that may improve coherence, and notes coherence is positively correlated with accuracy across models.
2 Coherence and Dutch Books
The paper operationalizes probabilistic coherence as resistance to Dutch books: forecasts are coherent exactly when no portfolio of bets guarantees a positive payoff. A linear-program measure quantifies the guaranteed arbitrage profit and depends on the queried event set.
- Known logical relationships among events partition outcomes into atoms, represented by an incidence matrix mapping atoms to events.
- A bet on event E_i with stake b_i costs b_ip_i and pays b_i when the event occurs, producing payoff b_i(M_ij − p_i) on atom ω_j.
- A Dutch book is a portfolio whose net payoff is strictly positive for every atom, and arbitrage profit measures guaranteed profit per unit of total stake.
- The stake constraint caps the notional amount wagered and keeps the maximized profit finite.
- The arbitrage-profit measure equals the sup-norm distance from stated probabilities to the set of coherent forecasts and is zero exactly when forecasts are coherent.
- The measure concerns response coherence rather than necessarily true beliefs, differs from accuracy, and weakly increases as events are added.
3 Data and Experimental Design
The experiments elicit forecasts from fifteen language models over variance-normalized stock-return events using a shared stock-date sample and multiple prompt protocols. Analyses average linear-program coherence scores across repeated passes and use date-clustered bootstrap intervals.
- Sample: The sample covers 100 stock-days drawn from CRSP returns and Refinitiv headlines, with 98 distinct stocks and historical and future-return requirements.
- Models: Fifteen open- and closed-weight model families are evaluated with identical prompts, provider-default reasoning settings, temperature one when available, and an 8,192-token limit.
- Event panels: For horizons one through four, event panels classify variance-normalized daily returns into four bins and include logically expressible events.
- Elicitation: Baseline prompts provide each stock’s previous 60 normalized daily returns and ask one event per independent session without revealing stock identity.
- Elicitation: Variants change one prompt feature at a time, including stock identity, headlines, grouped events, forecasting guidance, historical bin frequencies, red herrings, and response averaging.
- Scoring: Each prompt normally receives five passes, with a separate linear-program solution per pass and the average arbitrage profit used as incoherence.
- Inference: Confidence intervals are generally 95 percent percentile bootstrap intervals clustered by date over 20,000 replicates.
4 Results
Language-model forecast incoherence is widespread and depends on both the logical structure of queried events and the elicitation protocol. Richer dependencies and several contextual changes increase arbitrage profit, while some multi-outcome and averaging strategies reduce it.
- 4.1 Language Model Forecasts Are Incoherent: 0.002067 mean profit per unit of gross stake was observed for GPT-OSS-120B, with at least 0.1 percent in 48 of 100 stock-days.The 95% confidence interval was [0.001471, 0.002737].
- 4.1 Language Model Forecasts Are Incoherent: At least 0.5 percent profit occurred in 15 stock-days and at least 1 percent in 6, exceeding the maximum mechanically generated by four-decimal rounding.Rounding can produce at most 5 × 10^-5 profit per unit of gross stake.
- 4.1 Language Model Forecasts Are Incoherent: Incoherence varies across model families, with larger variation in coherence than in accuracy.The evaluation reports mean arbitrage profit and mean Brier score across 15 models.
- 4.3 Questions With Logical Dependencies Generate Additional Incoherence: Linking multiple assets or days raises incoherence, and adding joint events increases arbitrage profit beyond marginal-event forecasts.Matched queries also associate richer logical dependence with lower coherence when the number of queried events is held fixed.
- 4.4 Incoherence Varies Across Elicitation Strategies: Coherence changes substantially with elicitation details: statistics reduce arbitrage profit, whereas instructions, company identity, and headlines increase it.These experiments hold next-day, single-asset events fixed while varying prompt content and procedure.
- 4.4 Incoherence Varies Across Elicitation Strategies: Asking for the full event panel nearly eliminates incoherence, averaging repeated answers yields modest gains, and irrelevant context often increases incoherence substantially.The largest irrelevant-context increases follow direct intuitive forecasts and positive emotion; 95 percent of five-pass answers are identical.
5 Discussion and Limitations
The paper finds that forecast coherence varies across language models and rises in more demanding, logically richer settings. The metric also supports evaluation without outcome labels, while the findings remain bounded by the stock-return environment and elicitation conditions.
- Discussion: Incoherence varies considerably across recently released language models and increases in more demanding settings, such as more days or more assets.For GPT-OSS-120B, prompt details can also change incoherence, including cases where additional information lowers it.
- Limitations: The findings may not generalize beyond forecasting variance-normalized stock returns to other domains.The authors state that the forecasting environment provides no direct evidence about cross-domain generalizability.
- Limitations: The study has not shown that coherence rankings across models remain invariant under different elicitation conditions.The elicitation experiments are based on GPT-OSS-120B.
- Discussion: The coherence metric does not require outcome labels, allowing evaluation when accuracy cannot be assessed, including unresolved events.The paper applies this approach to events that have not yet resolved.
- Discussion: Alternative post-training strategies may improve coherence by encouraging models to reason across related events.The authors discuss test-time computation across events, group-level reinforcement learning, and the Dutch Book linear program as a possible exact set-level reward.
A Duality
The duality result connects arbitrage profit with the distance of stated forecasts from the set of coherent forecasts. Minimax duality and the dual norm establish this equivalence.
- A Duality: The appendix derives the equivalence between arbitrage profit and the sup-norm distance to coherent forecasts.This is the dual characterization stated in Section 2.
- A Duality: For bets b and atom distribution π, expected payoff is b′(Mπ − p), and linearity makes the minimum over π attain a simplex vertex.The simplex represents distributions over the atoms.
- A Duality: The arbitrage profit is written as max ∥b∥1≤1 min π∈∆m−1 b′(Mπ − p).The expression maximizes worst-case payoff over bets in the ℓ1 unit ball.
- A Duality: Minimax allows exchanging the optimization order, while maximizing over the ℓ1 ball yields the dual norm ∥Mπ − p∥∞.This produces the sup-norm distance formulation.
B Prompts
The prompt design elicits probabilities for stock-return events with known logical structure across assets, horizons, information conditions, and contextual details. Long-horizon variants show that coherence can also be evaluated before outcomes resolve.
- Baseline prompts: Queries provide normalized return histories and ask about return bins, complements, conjunctions, and multi-day events.The prompts use fixed wording for complements and cross-day conjunctions.
- Two-asset prompts: Two-asset queries define both assets’ scales and ask separately about marginal bins or jointly about same-day outcomes.The displayed history aligns both assets by date.
- Information arms: Information arms append observed bin frequencies or instructions to use those frequencies as starting estimates.The statistics are computed from the displayed 60-day history.
- Grouped arms: Grouped arms request probabilities for randomized labeled event sets, including complement pairs, relation-poor pairs, and full panels.The full-panel arm lists all fourteen events in one query.
- Red herring arms: Red-herring arms insert fixed personal, investment, anxiety, or gambling-related notes before the event question.These details vary independently of the return history and event request.
- Unresolved outcomes: For 2-, 5-, and 10-year horizons, prompts ask for conditional probabilities on the first NYSE session after the anniversary, given stock survival.The target is a one-day return, not cumulative intervening performance.
- Unresolved outcomes: Mean incoherence was 0.0080 [0.0064, 0.0099] at 2 years, 0.0070 [0.0056, 0.0088] at 5 years, and 0.0077 [0.0060, 0.0096] at 10 years.The exercise demonstrates coherence evaluation for unresolved outcomes.
C Elicitation
The elicitation pipeline uses multiple providers for the evaluated models and accepts only structured, bounded probability responses. Invalid or failed responses are retried and categorized by failure type.
- Elicitation: Each model is accessed through one provider, with temperature 1.0 and an 8,192-token output limit where supported.Providers include Groq, DeepInfra, Amazon Bedrock, Google Vertex, and Azure.
- Elicitation: Responses are accepted only when valid JSON contains exactly the requested fields and probabilities in [0, 1].Errors, refusals, truncations, and malformed responses trigger retries.
- Elicitation: Across completed GPT-OSS-120B workhorse runs, 4,744 retry attempts occurred across 4,114 requests out of 345,500.The passage attributes retries to rate limits, provider or response errors, interrupted dispatches, connection errors, timeouts, and missing routing metadata.
D Linked Elicitation
The experiments compare linked and dispersed event sets while holding event-set size fixed, showing that added logical structure is associated with higher incoherence across both assets and days.
- Event-set construction: Linked event sets contain the joint events needed to decompose one marginal, whereas dispersed sets distribute joint events across different marginals.For both two-asset and two-day analyses, linked sets impose one additional equality between joint and marginal probabilities.
- Two assets: Two-asset panels include four marginal events and four joint events, with linked panels completing one marginal’s decomposition and dispersed panels applying a fixed cyclic shift.The construction produces eight linked and eight dispersed event sets when each asset is used as the base.
- Results: 0.00097 [0.00081, 0.0012] higher incoherence was observed for linked two-asset event sets than for dispersed sets.The number of elicited events was held fixed while relationship strength varied.
- Results: 0.0021 [0.0019, 0.0023] higher incoherence was observed for linked two-day event sets than for dispersed sets.This experiment likewise held the number of elicited events fixed while varying relationships across events.
- Two days: Two-day panels use four marginal events and candidate joint events from same-bin and different-bin branches, with linked and dispersed panels selecting these branches differently.Dispersed panels make one branch selection per bin, while linked panels complete one bin, omit another, and select one branch for the remainder.