Source-linked AI summary
LLM-Derived Preference Judgments Are Not Self-Consistent
Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
TL;DR
LLM-based preference learning assumes numerical judgments can be represented by one stable utility function, but their self-consistency remains uncertain. This paper audits that assumption across query types, domains, and models, finding persistent inconsistencies that sometimes reverse choices.
Problem
The paper asks whether numerical LLM preference judgments can be reproduced by one stable, quasi-linear, dollar-denominated utility.
Method
The authors develop an audit protocol that fits unrestricted item values and jointly tests whether query means are self-consistent.
Results
Across all six models, the joint self-consistency claim is rejected; P2 disagreements are frequent, often substantial, and sometimes reverse preferred offers.
Takeaways & Limitations
LLM preference judgments depend on query type and cannot be treated as interchangeable under a shared quasi-linear dollar utility.
Takeaways & Limitations
The study audits internal coherence rather than human fidelity, uses no human data, and its controlled setup is not representative of real users or interactions.
Abstract
from arXiv · showhide
Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.
1 Introduction
The paper audits whether numerical LLM preference judgments from different monetary query types can be reproduced by one stable quasi-linear, dollar-denominated utility. It develops joint and local consistency tests and applies them across three domains and six models, finding persistent disagreement across query types.
- Motivation: Preference-learning systems commonly rely on a stable latent utility function to represent or infer preferences from LLM judgments.Examples include direct utility fitting, latent item-utility beliefs, and transitivity- and completeness-based choice models.
- Research question: The central question is whether population means of numerical LLM judgments can be reproduced by one stable quasi-linear, dollar-denominated utility.The paper distinguishes downstream decision performance from the more basic measurement question of mutual consistency across prompts.
- Measurement setup: Item queries elicit maximum willingness to pay, whereas offer-pair queries elicit the signed price change that would make the user indifferent between two offers.The sign encodes ordinal direction, while the magnitude is a cardinal monetary judgment.
- Measurement setup: The audit assigns each text-described item an unrestricted dollar utility and imposes only quasi-linearity in listed price, without feature-based utility assumptions.Listed prices provide the common cardinal scale, while item descriptions remain atomic alternatives.
- Audit and findings: The protocol combines a joint bootstrap test of shared utility fit with local P1 and P2 diagnostics, and experiments across three domains and six models find persistent query-type disagreement.P1 tests path additivity among offer-pair queries; P2 compares an offer-pair estimate with the price-adjusted difference between two item-query estimates.
2 Related Work
Prior work uses LLM-elicited preference signals in downstream selection and optimization, with varied feedback formats and utility representations. This paper instead fixes a monetary scale, combines item-query and offer-pair responses, and audits their agreement under quasi-linearity before fitting utilities.
- LLM-derived and LLM-mediated preference signals: LISTEN, LILO, and PEBOL use LLM-derived preference signals as inputs to selection, utility fitting, or optimization procedures.LISTEN maps natural-language preferences to utilities, batch choices, or selected items; LILO fits Gaussian-process utilities from scalar or pairwise feedback.
- Feedback form and shared scalar representation: LILO’s scalar-versus-binary comparison changes the response scale, observation model, and sequentially selected points simultaneously.Related LLM feedback and evaluation work also uses bounded or rubric-anchored pointwise scores and pairwise judgments.
- Feedback form and shared scalar representation: This paper expresses item-query and offer-pair responses in dollars, holds the monetary scale and compared items fixed, and tests their agreement under quasi-linearity before downstream fitting.Listed price supplies the external numeraire, and audited items receive unrestricted utilities rather than a feature-based model.
- Utility representation and consistency audits: The consistency assumption builds on utility representation of complete and transitive preferences, while the paper’s diagnostics also relate to HodgeRank’s global potentials and cyclic inconsistency.The setting combines item-query anchors with offer-pair measurements and tests whether they agree under quasi-linearity.
3 Utilities and LLM Estimates
The paper models LLM preference responses as dollar-denominated utility judgments under quasi-linearity in money, while allowing item utilities to be unrestricted. It tests whether one utility function explains all query means and measures persistent discrepancies when it does not.
- Utility model: The model assumes quasi-linearity in money but imposes no linearity, smoothness, separability, or additivity on item utilities.Items are represented by textual descriptions, including flights, apartment rentals, and hotel rooms.
- Query types: Item queries estimate U(x) relative to not buying, whereas offer-pair queries estimate the signed price adjustment needed to make two offers indifferent.Positive adjustments mean the target offer could be more expensive; negative adjustments mean it must be cheaper.
- Self-consistency hypothesis: Self-consistency requires a single utility U to reproduce the population mean of every audited query.The alternative is that every candidate utility leaves at least one query mean unexplained.
- Best-fitting utility and discrepancies: The best-fitting utility U⋆ minimizes average squared discrepancies across queries; under self-consistency every discrepancy is zero, whereas the alternative leaves persistent mean disagreement.Repeated calls reduce sampling variation but do not remove persistent discrepancy δ(c).
- Local audits: Local tests compare alternative estimates of the same utility difference, and rejecting a local null rejects self-consistency for those queries but does not establish global inconsistency when it is not rejected.Comparisons include direct offer-pair estimates versus path sums and versus item-query differences.
- Global audit: T is the smallest RMSE achievable by assigning one free utility value to every audited item, with every prespecified query weighted equally.As repetitions grow, T approaches zero under self-consistency and the RMSE of persistent discrepancies under the alternative.
4 Audit Protocol
The audit evaluates cardinal preference consistency across controlled flight, apartment, and hotel item sets using stateless prompts and separate model-domain audit groups. It prespecifies comparison templates, repeated sampling, multiplicity correction, and distinct global and local residual scales.
- Audit design: The audit constructs interpretable item sets for flights, apartments, and hotels, while using only item identity, listed price, and query incidence in the audit itself.Flight and apartment descriptions vary continuous, ordinal, and categorical features; hotels form nine offers from three room items and three prices.
- Audit design: Each model has nine audit groups formed by crossing one domain, utterance, item set, and query collection, with utility fits and test statistics computed separately.Three fixed utterances per domain vary emphasis on price and domain-specific features.
- Comparison design: The prespecified design uses 16 endpoint-comparison templates under three domain-specific utterances, producing 48 comparison instances across nine audit groups.The templates cover six flight, six apartment, and four hotel comparisons.
- Query protocol: The audit uses independent, stateless prompts with one canonical wording per query type and no earlier path questions or answers.Across the design, 60 item queries support 96 P1 and 48 P2 comparisons, though pooled residuals are not independent.
- Models and inference: Six models receive 15 calls per query, and group-level rejections use Bonferroni-adjusted p ≤0.0056 to control any false rejection probability at 5%.The audited models are Claude Opus 4.8, Gemini 3.5 Flash, GPT-5.5, GPT-OSS 120B, Llama 3.3 70B, and Qwen 3.6 27B.
- Measures: Global Tg is reported in dollars relative to mean endpoint listed prices, whereas local residuals use each pair’s mean endpoint price and exclude intermediate path prices.A 0.10 local residual represents disagreement equal to 10% of the mean endpoint price; ordinal reversals and local rejections require confidence intervals excluding zero.
5 Audit Results
The audit rejects exact self-consistency across models and shows that the best-fitting unrestricted utilities still leave substantial residual disagreement, especially for hotels and P2 comparisons. Local results indicate that query construction affects estimated utility differences, while P1 failures are less consistent across models.
- Global audit: Every model rejects the joint claim that all nine audit groups are self-consistent after Bonferroni correction; five reject all nine groups, while Qwen rejects four.Because each audited item receives an unrestricted utility, the rejections are not failures of a particular feature-based or parametric utility model.
- Global audit: 18.9–44.9% RMSE occurs for hotels versus 1.6–6.1% for flights and apartments after fitting the best possible utility.These are descriptive rather than controlled domain comparisons because query geometry and fit rank differ by domain.
- Local failures: 41.7–87.5% of P2 residual confidence intervals exclude zero, and supported reversals occur for 2.1–12.5% of 48 prespecified comparisons per model.Each supported reversal makes a two-offer selector based on item-query means choose the opposite offer from one based on the offer-pair mean.
- Local failures: P1 rejection rates range from 6.2% to 68.8% across models; supported reversals are rare, and price-normalized magnitudes are generally smaller than for P2.The results oppose unrestricted composition of offer-pair estimates in the audited design but do not establish a robust cross-model ordinal effect.
- Stress test: Mean price-normalized P1 error is higher at k = 8 and k = 16 than at k = 2 for every model in the supplementary stress test.The four-path design is small and nonmonotone, so it does not establish a general scaling law.
6 Implications for Preference Learning
LLM-derived preference judgments are not necessarily unusable, but their dependence on question framing and model choice requires careful auditing and, where possible, comparison with human judgments. Self-consistency can support preference learning, yet fidelity to human values requires additional validation.
- Implications for Preference Learning: Algorithm designers should treat LLM preference information cautiously because judgments depend on question wording and the choice of LLM.When possible, they should corroborate LLM judgments with real human judgments.
- Implications for Preference Learning: Audits with new items and price changes can test consistency before numerical elicitation is deployed.Metrics from these audits can guide selection of an LLM with better self-consistency.
- Implications for Preference Learning: Using an LLM with poor self-consistency risks results that depend arbitrarily on the specific form of questions asked.The paper proposes using its self-consistency metrics to inform model choice after auditing an application domain.
- Implications for Preference Learning: Self-consistency is necessary but not sufficient for fidelity, because a single quasi-linear dollar utility may still assign values a person would not endorse.The audit design cannot detect this gap without eliciting the same item and offer-pair queries from the people who supplied the preference descriptions.
- Implications for Preference Learning: Human-subject research can compare query constructions and calibrate LLM inconsistencies against human inconsistencies, including documented willingness-to-pay framing effects.Eliciting matched queries from people supplying the preference description could identify which construction better recovers stated human values.
7 Conclusion
Across all six models, the joint claim that all nine audit groups are self-consistent is rejected after Bonferroni correction. P2 disagreements are frequent and sometimes reverse preferred offers, while P1 and path-length findings are more model- and design-dependent.
- Conclusion: 9 audit groups’ joint self-consistency is rejected across all six models after Bonferroni correction.The audit tests whether heterogeneous numerical LLM answers estimate one utility.
- Conclusion: P2 disagreements are frequent, often large relative to price, and sometimes reverse the preferred offer.These disagreements identify substantial departures from self-consistent judgments.
- Conclusion: P1 and path-length results caution against unrestricted composition but depend more on the model and design.The audit does not establish human fidelity; it identifies when query types or paths produce inconsistencies.
Limitations · A Formal basis of the audit
The paper’s audit tests internal coherence rather than human fidelity, under restrictive experimental and modeling choices. Its formal basis normalizes item utilities and shows that self-consistency implies zero P1 and P2 residuals, with nonzero discrepancies rejecting coherence.
- Limitations: The study audits internal coherence rather than human fidelity and uses no human data.Its controlled items, three domains, six models, fixed prompts and provider versions, stateless calls, and 15 completions may not represent real use.
- Limitations: The experimental design may differ from real users, larger rankings, paraphrases, model versions, and conversational histories.The quasi-linear dollar model also excludes wealth effects, binding budgets, and cases without finite compensation; P2 covers price-free item queries only.
- A.1 Why the queries estimate item utilities and offer-utility differences: The measurement model subtracts buying-nothing utility as a zero reference, although this normalization is not stated in the LLM prompt.Under this normalization, U(x) denotes the utility gain from obtaining x relative to buying nothing.
- A.1 Why the queries estimate item utilities and offer-utility differences: Direct queries target item utility U(x), while offer-pair queries target signed target-minus-source utility differences between offers.Self-consistency therefore requires E[Y (x)] = U(x).
- A.2 Why P1 and P2 follow from self-consistency: Under self-consistency, population P1 and P2 residuals are zero.These residuals compare path-based offer differences and item-query levels with the corresponding utility representation.
- A.2 Why P1 and P2 follow from self-consistency: A nonzero expected P1 or P2 discrepancy is sufficient to reject self-consistency on those queries.The proof uses utility-difference expectations and telescoping sums along paths.
- A.2 Why P1 and P2 follow from self-consistency: Offer-pair estimates admit one offer utility exactly on each connected component when every cycle sum is zero.Zero cycle sums make path-defined utilities independent of the chosen reference path; utility differences also imply antisymmetry and transitivity of signs.
B Experimental reproducibility details … B.4 Model and API configuration
The audit uses a controlled multi-domain design with fixed prompts, repeated API calls, and documented analysis settings. Its reproducibility details specify query construction, model configuration, pilot-freezing procedures, and computational environment.
- B.1 Domains, items, and query counts: The audit covers flights, apartments, and hotels using source offers and controlled item features, while listed price serves as the monetary numeraire.Item features construct the controlled design but are not covariates in the utility fit.
- B.2 Audited preference utterances: Preference utterances vary by domain and focus on price, time, space, commute, value, or quality tradeoffs.The audited utterances include three preference variants for flights, apartments, and hotels.
- B.1 Domains, items, and query counts: Across domains, 300 queries with 15 completions each produce 4,500 planned calls per model.Each source-target pair contributes one direct offer-pair query, four path-step queries, and two P1 comparisons.
- B.3 Prompts and exact query specification: The released query file records concrete system and user prompts for all 300 queries, with canonical wording within each query type.Only the preference utterance and item or offer content vary as specified.
- B.4 Model and API configuration: The API runner requested temperature 1 and seeds 10000 through 10014 where supported, with a nominal 2,000-token response limit.GPT-5.5 instead used an 8,000-token completion limit; limits were upper bounds and did not match reasoning budgets across providers.
- B.4 Model and API configuration: Each request asked for one response; provider-default generation controls were preserved, and bootstrap analyses used B = 2000 and B = 5000 draws with separate seeds.Stop sequences and log probabilities were disabled, while Groq explicitly used stream=false and other clients used nonstreaming calls.
- B.4 Model and API configuration: Final prompts, temperature, 15-call sample size, seed schedule, reasoning settings, bootstrap sizes, and testing level were fixed before corresponding full-run analyses.Separate pilot responses were excluded from reported audit results; GPT-OSS pilots reached finite-parse rates of 67.8%, 93.3%, and 100% at ceilings of 400, 1,200, and 2,000 tokens.
- B.4 Model and API configuration: Inference used hosted APIs rather than local accelerators and was orchestrated on an Apple M5 MacBook Pro running macOS 26.4.1 and Python 3.12.6.The environment used 10 CPU cores, 16 GB memory, arm64, and specified package versions for analysis and API clients.
C Supplementary nonfinite-output diagnostics · D Supplementary path-length stress check
The supplementary diagnostics found nearly universal finite parseability, with one GPT-OSS empty response and two unambiguous Gemini parses involving trailing underscores. The path-length stress check found higher mean P1 disagreement at k = 8 and k = 16 than at k = 2 for every model, without establishing a general scaling law.
- C Supplementary nonfinite-output diagnostics: All responses from Llama 3.3, Qwen 3.6, GPT-5.5, Opus 4.8, and Gemini 3.5 Flash yielded finite parsed values.
- C Supplementary nonfinite-output diagnostics: GPT-OSS produced one empty offer-pair response.
- C Supplementary nonfinite-output diagnostics: Two Gemini responses contained trailing underscores but were unambiguously parsed as 30.
- D Supplementary path-length stress check: The path-length stress check used four prespecified smooth single-feature paths: two apartment square-footage paths and two flight travel-time paths under one canonical utterance.
- D Supplementary path-length stress check: Mean absolute P1 residual | bR1,k| was normalized by the fixed mean price of the endpoint offers, sab = (p+q)/2.
- D Supplementary path-length stress check: Bootstrap 95% intervals were computed over the four paths, and the stress test did not establish a general scaling law.
- D Supplementary path-length stress check: For every model, mean error at k = 8 and k = 16 exceeded that at k = 2, although the curves were not all monotone.
- D Supplementary path-length stress check: 1,860 calls per model in the v2 experiment produced 11,160 responses, all with finite parsed values.
E P2 results by domain · F Statistical details · F.1 Audit-group self-consistency tests
P2 disagreements persist across all three domains and every model–domain combination, while the audit tests self-consistency within nine domain–utterance groups using a fitted-null bootstrap and Bonferroni correction. Domain counts are descriptive rather than population estimates, and utilities are conditional on each group’s utterance and item set.
- E P2 results by domain: Domain counts are descriptive because the audit was not designed to estimate population differences between domains.
- F.1 Audit-group self-consistency tests: Nine audit groups combine three domains with three selected utterances per domain, and each group has its own utterance-conditional utility over that domain’s item set.Utilities are neither shared nor directly comparable across groups.
- F Statistical details: Increasing completion counts reduces the sampling-variation component of the expected squared statistic but not the component from discrepancies no utility assignment can absorb.The calculation interprets the finite-sample baseline; the bootstrap test uses its own reference distribution.
- F Statistical details: The within-query nonparametric bootstrap preserves each query’s empirical response distribution, accommodating non-Gaussian responses and query-specific variances.Synthetic query means are generated under the fitted self-consistency null, then the utility is refit and RMSE recomputed in every draw.
- F.1 Audit-group self-consistency tests: The observed RMSE measures disagreement among original query means after fitting the utility, whereas bootstrap RMSEs reproduce sampling variation absorbed by refitting.Centered completion pools retain each query’s observed variability while imposing mean-zero noise under the fitted null.
- E P2 results by domain: Supported reversals occur in all three domains, and every model–domain combination contains supported cardinal disagreement.A supported reversal requires offer-pair and item-query 95% intervals to exclude zero in opposite directions; the final column counts cases with disagreement lower bounds of at least 25% of mean listed price.
- F.1 Audit-group self-consistency tests: Bonferroni testing rejects a group only when p ≤0.05/9 = 0.0056, controlling the chance of falsely rejecting any self-consistent group at at most 5%.The correction is needed because the model-level claim requires all nine audit groups to be self-consistent.
F.2 Local residual uncertainty
Local residual uncertainty is estimated from query-level sampling variation, with residuals combining independent query averages and fixed listed-price adjustments. Confidence intervals use normal approximations, while lower-bound curves are comparison-specific rather than simultaneous bands.
- Query-level uncertainty: Each query average has standard error s_c/√n_c, where n_c is the number of finite completions and s_c their sample standard deviation.Queries without finite responses are excluded from residual analyses and counted in Table 6.
- Residual uncertainty: Each P1 or P2 residual is a signed sum of three query averages whose constituent queries use disjoint LLM calls.The signs do not affect the variance calculation, and fixed listed-price adjustments contribute no sampling variance.
- Intervals and thresholds: 95% intervals are Wald intervals, b_R ± 1.96 c SE(b_R), and local rejection occurs when the interval excludes zero.Threshold curves divide |b_R| by the fixed mean-price scale s_ab = (p + q)/2, so the standard error scales by the same fixed denominator.
- Intervals and thresholds: Dashed curves show comparison-specific nonnegative magnitude lower bounds and are not simultaneous confidence bands.Thus, the curves do not provide a joint confidence guarantee across comparisons.
- Independence assumptions: Repeated completions within fixed model–query pairs are treated as independent and identically distributed, with variances allowed to differ across queries.Queries use separate prompts and calls, but paths sharing a source-target pair share a direct estimate, inducing positive residual correlation that is not modeled.