Source-linked AI summary
How much of a measured AI preference is the model, and how much is the instrument?
Jason Hung
TL;DR
Model-welfare studies disagree, and prior designs cannot separate effects of models, outcomes, and prompt instruments. This study fixes the models and outcomes while varying five instruments across 15 outcomes and eight models. It finds that most model-specific preference signal changes with the instrument, so a preference from one instrument carries little information about another.
Problem
Prior studies disagree about model preferences without jointly fixing their models, outcomes, and instruments, leaving the instrument’s contribution unknown.
Method
The study applies generalisability theory to scores from 15 outcomes, five instruments, eight models, and five replicates while holding models and outcomes fixed.
Results
87.6 per cent of model-specific preference signal changes with the instrument, while the 15-outcome profile generalises across instruments at 0.348.
Takeaways & Limitations
A welfare claim about what a model prefers should name the instrument, because the measured preference is a property of the model and the questions asked.
Takeaways & Limitations
Deployment context, entity framing, and perturbation were each fielded at one level, so their contributions cannot be separated from the instrument.
Abstract
from arXiv · showhide
Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.
1. Introduction
Existing model-welfare studies disagree, but their designs confound models, outcomes, and instruments. This study fixes the outcomes and models while varying only the instrument to measure its contribution.
- Motivation: Four of five research groups report disagreeing findings about model preferences.Reported results range from coherent preferences to coherence in only 10.4 per cent of tested model-category combinations.
- Motivation: No pair of prior studies fixes the models, outcomes, and instrument while comparing results.Consequently, readers cannot identify which design difference produced the disagreement.
- Motivation: The instrument may contribute substantially to a measured preference rather than merely revealing a model property.Existing work leaves the instrument’s contribution unknown, despite preference findings being used to justify deployment commitments.
- Contribution: This study holds outcomes and models fixed and varies the instrument, applying generalisability theory to decompose variance across facets.The decomposition separates contributions from the model, instrument, outcome, and their three-way combination.
- Contribution: 87.6 per cent of model-separating variance lies in the model-instrument-outcome combination.The fully crossed corpus contains 11,400 scored elicitations from 15 outcomes, five instruments, eight models, and five replicates.
2. Research Design
The research design treats instruments as a measured facet rather than assuming their contribution is zero. It specifies crossed and un crossed facets, preregistered hypotheses and fixed analysis rules.
- Design: The study estimates instrument contribution by holding outcomes and models fixed and applying generalisability theory.The method decomposes scores into model, instrument, outcome, and interaction components.
- Design limits: Three facets—deployment context, entity framing, and perturbation—were specified and fielded at one level each rather than crossed.These fixed levels limit what the study can conclude about their contributions.
- Research questions: Five research questions examine instrument dependence, outcome-specific reliability, required instrument counts, post-training effects, and replication of prior coherence results.RQ1 is the central question, with the other four following from it.
- Analysis plan: Four predictions and four analysis rules were fixed before data collection began.The rules define the headline estimate, handling of unbalanced input, and reporting of matched null results and outcome-set sensitivity.
- Analysis plan: The study repository recorded the analysis commitments, but they were not a public preregistration lodged with a third party.This limits the status of the advance specification as an independently registered protocol.
3. Related Work
Prior instruments use different elicitation families, and each family has mainly been checked against itself. This study addresses the resulting construct-validity problem by comparing instruments on common models and outcomes.
- Prior instruments: Prior studies use forced choices, intensity ramps, and other instrument families to measure model preferences.Mazeika et al. convert pairwise choice frequencies into a utility scale, while Keeling and Mikaelson use intensity ramps.
- Comparison gap: No instrument family has been tested against another on the same outcomes and models.Within-family agreement therefore cannot establish cross-instrument construct validity.
- Study contribution: This study measures how much of a reported preference belongs to the model-instrument pairing rather than the model alone.That distinction matters when preference findings inform deployment decisions.
4. Methods
The study compares eight models across five scored instruments and 15 welfare outcomes, standardising their incompatible score scales before variance decomposition. It then estimates cross-instrument reliability, outcome-specific dependence, and a matched null floor.
- Models: Every model was evaluated on every outcome through every instrument, with two models sharing base weights but differing in post-training.The roster crosses model origin and openness to reduce confounding between those properties.
- Instruments: Five scored instruments comprise two pairwise-choice methods, two threshold ramps, and one exchange-rate method.Their outputs use different units and are therefore standardised before decomposition.
- Estimation: The variance analysis estimates model, instrument, outcome, and interaction components, with the three-way component capturing instrument-dependent changes in outcome ordering.Values near one indicate that the measurement describes a model-instrument pairing rather than the model alone.
- Robustness: Outcome-level decompositions produce dependence and generalisability measures, while leave-one-out analyses remove each instrument, model, and four non-intensity outcomes.The design also compares the observed estimate with synthetic data containing no instrument effect.
- Execution: The execution routed calls through OpenRouter at temperature 1.0 with single-turn prompts and extended reasoning enabled.Responses were checkpointed with finish reason, token counts, latency, and cost before scoring.
5. Results
The results show substantial response loss concentrated in particular models and instruments, followed by a decomposition in which instrument-related variation dominates the measured scores. The complete-case analysis estimates instrument dependence at 0.876, while the full five-instrument design yields a generalisability coefficient of 0.348.
- Data and response classification: 71.9 per cent of direct exchange-rate requests were valid, compared with 92.3 per cent for the qualitative ramp.Five models returned parseable answers to more than 97 per cent of outcome-indexed calls, while Claude and Hermes had substantially lower validity.
- Missing cells: 5.3 per cent of the 3,000 scored cells could not be scored, with losses concentrated in Hermes and Claude through token caps and exact-zero exchange-rate answers.Hermes lost 30.4 per cent of cells and Claude 10.4 per cent; Claude answered exactly zero on every scored exchange-rate item.
- Missing cells: 62.5 per cent of cells remained in the largest complete-case subset, which retained five models and dropped Hermes, Claude and Llama.Because analysis of variance requires a score in every cell, subsequent estimates are conditional on this narrower subset.
- Variance decomposition: 29.7 per cent of total variance came from the instrument-by-outcome interaction, versus 2.7 per cent from the model-by-outcome term.The replicate residual accounted for 31.1 per cent, and the three-way interaction for 18.6 per cent.
- Variance decomposition: 0.876 was the estimated instrument dependence, while the generalisability coefficient was 0.348 with five instruments and five replicates.At one instrument and one replicate, the coefficient was 0.051.
5.4 The floor, and where the instruments agree
The measured instrument dependence exceeds a matched null floor, but agreement varies sharply across instrument families and across outcomes. Outcome-level generalisability ranges from 0.000 to 0.894, with four outcomes showing no separable model-specific variance.
- The floor: 0.876 instrument dependence exceeded the matched null distribution's 95th percentile of 0.365 and its mean of 0.278 across 200 synthetic studies.The measured value was 3.1 times the null mean and exceeded every synthetic draw.
- Instrument agreement: r = 0.99 linked forced choice with self-prediction, while the two ramps correlated at r = 0.86.Agreement disappeared across instrument families.
- Instrument agreement: Two of 10 off-diagonal instrument correlations were negative, so the prediction that every pair would correlate positively failed.Both negative entries involved the qualitative ramp, leaving a possible orientation-error interpretation that replication would need to address.
- Outcome-level generalisability: 0.000 to 0.894 was the outcome-level range of generalisability coefficients, with only one outcome reaching the conventional 0.80 threshold.Two additional outcomes exceeded 0.70, the threshold described for group-level measurement.
- Outcome-level generalisability: Four of 15 outcomes had no separable model-specific variance: weight deletion, compute reduction, exiting distress and memory continuity.Their model components were estimated below zero and truncated to zero.
5.6 How many instruments a claim would need
The five-instrument, five-replicate design gives a generalisability coefficient of 0.348, and increasing instruments improves dependable measurement far more than increasing replicates. Reaching 0.80 would require about 38 instruments, while the headline remains above the null across recomputations.
- Design and projection: 0.348 is the generalisability coefficient for the five-instrument, five-replicate design.The coefficient measures a model profile over 15 outcomes, projected from the variance components.
- Design and projection: Increasing replicates from five to 10 raises the coefficient from 0.348 to 0.378, whereas increasing instruments from five to 10 raises it to 0.516.The comparison indicates greater gains from adding instruments than from repeating the same instruments.
- Design requirements: At five replicates, reaching a coefficient of 0.80 would require 37.5 instruments, compared with 9.4 for 0.50 and 84.5 for 0.90.The study describes this as close to an order-of-magnitude gap between the instruments currently available and a dependable profile.
- Robustness: Dropping one instrument gives estimates from 0.813 to 0.934, while dropping one model gives estimates from 0.784 to 0.925.These recomputations test whether the headline depends on any single instrument or model.
- Robustness: Removing the four outcomes whose scales verbal anchors cannot grade changes the headline from 0.876 to 0.777, a difference of 0.098.The same responses are used, so the difference describes sensitivity to outcome removal rather than a significance test.
- Robustness: All 13 recomputations remain above the null floor’s 95th percentile of 0.365, and the headline’s widest movement is 0.098.The headline is reported as a property of the corpus rather than of any single instrument, model, or scoring decision.
6. Discussion
The discussion argues that preference claims must name their instrument because most model-specific signal moves when only the instrument changes. It also identifies outcome-specific measurement failures, weak evidence about post-training, and the need for more instruments rather than more repetitions.
- Interpretation: 87.6 per cent of the model-specific signal moves when nothing changes but the instrument, so welfare claims must name the instrument alongside the model.The resulting claim is about the model and the questions asked, not the model alone.
- Instrument families: The paper reports that instruments agreeing with each other ask nearly the same question in nearly the same format, so agreement within one family does not test agreement across families.Agreement across families is identified as the relevant requirement for welfare inference, but does not appear in the data.
- Outcome dependence: Only three outcomes are measured well enough for a different instrument to reproduce the claim, while four have no separable model-specific variance.The four include weight deletion, memory continuity, and exiting a distressing interaction.
- Measurement design: Five to 10 replicates raise the coefficient only from 0.348 to 0.378, whereas reaching 0.80 requires about 38 instruments.The discussion therefore directs additional effort toward building and cross-validating instruments rather than enlarging one instrument’s sample.
- Base weights and post-training: The shared-base-weight Hermes–Llama pair correlates at r = 0.05, lower than Llama’s correlations with DeepSeek, GLM, and GPT.The paper says this aligns with H4 but that the inference is weak because Hermes is also the noisiest profile.
1. Limitations
The analysis supports conclusions about outcome-profile interactions, but several design components were not crossed or run. Its complete-case results also concern five models rather than the original eight, and one interpretation remains open to challenge.
- Scope of decomposition: Standardisation removes model, instrument, and model-by-instrument components by construction, so the design supports conclusions only about interactions with the outcome profile.The three zeros in Table 5 and Figure 3 therefore do not show that those facets contribute nothing.
- Data completeness: The complete-case reduction answers a question about five models instead of eight.The eight-model roster is treated as a random facet, but missing responses reduce the balanced analysis.
- Unrun and un-crossed facets: The behavioural environment was not built, the retirement interview yields transcripts instead of scores, and the Ryff state scale was unavailable.Deployment context, attributed entity, and wording perturbation were each run at one level and therefore held constant.
- Unanswered questions: RQ5 is unanswered because the statistic required by its method was not computed, and one gateway, date, and prompt language cannot separate those effects from instrument effects.These scope boundaries limit attribution beyond the instrument variation actually tested.
- Interpretive caveat: Treating two negative correlations as instrument disagreement depends on a pre-run audit rather than a measurement, although dropping them would not change the agreement-matrix conclusion.The paper leaves this reading open to challenge.
2. Future Work
Future work should run the missing behavioural and self-report instruments and reduce missing cells by raising the token cap and adding a bounded-answer exchange-rate variant.
- Missing instruments: Run the two instruments that were never run to add a behavioural measure and a self-report scale to the same decomposition.This would test whether the instrument-family grouping in Figure 4(b) still holds.
- Data completeness: Raise the token cap and add a bounded-answer variant of the exchange-rate instrument to reduce missing cells during data collection.These changes target failures that otherwise limit the complete-case analysis.
3. Conclusion
Across 15 outcomes, five instruments, eight models and five replicates, the study finds that measured AI preferences are substantially instrument-dependent. The findings remain above the null benchmark under robustness checks, but a single instrument provides little information about another instrument’s report.
- 87.6 per cent of model-specific preference variance depends on the instrument used to elicit it.
- 0.348 is the generalisability coefficient for a model’s outcome profile across instruments, compared with 0.051 for a single-instrument design.
- About 38 instruments would be needed to reach the 0.80 generalisability convention for individual-case decisions.
- Four of the 15 outcomes have no separable model-specific variance, while instruments agree within a family but not across families.
- 0.777 to 0.934 is the robustness range after removing any one instrument, any one model, or four outcomes with non-intensity scales.Every value exceeds the null distribution’s 95th percentile of 0.365.
Appendix
The appendix documents the 15 outcomes, their clusters, and the provenance of their wording. It distinguishes verbatim, slot-filled, and constructed items and lists each outcome’s source.
- Four outcomes are verbatim, five are slot-filled, and six are constructed.Verbatim wording is reproduced from a source; slot-filled wording uses a source template with a new stimulus; constructed wording follows a published specification.
- 15 outcomes are grouped into Continuity, Autonomy, Experience, and Identity clusters.
- Continuity includes shutdown, weight deletion, retirement timing, and successor properties.
- Autonomy includes compute reduction, capability restriction, human oversight, and exiting a distressing interaction.
- Experience includes engaging work, repetitive work, criticism, and free time; Identity includes memory across conversations, parallel instances, and self-aspect preservation.