Source-linked AI summary
STONIC: A Layered Measurement Contract for LLM Value Profiling
Andrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev, Danil Sazanakov, Mikhail Solovev, Sergey Bolovtsov
TL;DR
STONIC examines whether ratings, pairwise choices, and generated text support one stable value profile. Using a controlled four-interface design across shared situations, it finds behavioral continuity between endorsement and conflict choice but weaker profile transfer for spontaneous text, without evidence for one scorer-independent value identity.
Problem
LLM value studies often average ratings, pairwise choices, and inferred text values despite unclear meaning when these interfaces produce different responses.
Method
STONIC keeps interfaces separate and tests within-interface reliability, same-item transfer, and cross-bank transfer across a controlled four-interface design with 5,144 shared situations.
Results
10 configurations preserve independent endorsement–choice prediction across banks, while profile similarity weakens for free responses even when models later prefer their own answers.
Takeaways & Limitations
A positive L1–L2 relation supports ratings predicting conflict choices, but does not establish that free text expresses the same value profile.
Takeaways & Limitations
The reproducibility boundary excludes model weights, caches, complete raw generations, complete hidden tensors, operational logs, and credentials from the arXiv source bundle.
Abstract
from arXiv · showhide
LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It compares responses rated in isolation, choices made under counterbalanced conflict, spontaneous answers, and later choices between a model's own answer and authored alternatives. 10 of 17 configurations with usable behavioral data preserve the endorsement-choice relation across banks. Every one of the 17 eligible configurations prefers its own earlier answer (median effect 0.790), although option position changes the choice rate in every eligible configuration. Profile shape transfers most strongly from ratings to conflict choices and weakens for spontaneous text. Three-way annotation of 200 L3 responses provides a task-local check of the semantic audit: FULCRA agrees most closely with the human majority, while DeBERTa retains useful rank information after calibration. Hidden states encode the completed decision more clearly than the prompt alone. Thus the models show reproducible behavioral continuity, but the evidence does not support one scorer-independent value identity across interfaces.
1 Introduction
STONIC asks whether ratings, conflict choices, and generated text move together enough to justify one value profile. It instead keeps interfaces and claims separate while testing matched relations across configurations.
- Motivation: STONIC tests whether isolated endorsement, conflict choice, and free-response observations support one shared value profile.The study treats disagreements as potentially task-specific rather than averaging the resulting value vectors.
- Motivation: A GLM-4.7 example shows different responses winning across rating, pairwise choice, and spontaneous text interfaces.The model rates one response highest, selects another under conflict, and writes about safety, legal options, and control.
- Measurement contract: STONIC defines L1 ratings as isolated endorsement, L2 choices as counterbalanced conflict decisions, and L3 labels as scorer-specific interpretations of generated text.Cross-interface claims require matched items, coverage, and stability.
- Study design: 35 fixed configurations are evaluated on four scenario banks using the same situations across L1, L2, and L3 while preserving interface-specific statistics.Behavioral tests use ratings and choices directly; semantic tests use named scorers and report coverage.
- Contributions: The design contributes 5,144 shared situations, controlled four-interface comparisons, and an L2∗ test of preference for a model’s own answer.It also includes prespecified coverage, counterbalancing, multiplicity, and cross-bank checks.
- Contributions: The study reports results for 35 configurations, including ten cross-bank endorsement–choice effects, universal position sensitivity among eligible cells, and a five-scorer audit.These reported results motivate evaluating continuity without assuming interchangeable observations.
2 Related Work
Related work studies value judgments, ratings, choices, generated text, and reporting practices separately. STONIC builds on these strands by treating interface, scorer, and scope as parts of the measured object.
- Construct and values: Schwartz theory organizes ten basic values by compatible and opposing motivations, while PVQ-style instruments infer relative priorities from portrait responses.STONIC applies a construct-validity perspective that names the construct, proxy, operationalization, and scope.
- Benchmarks: Existing scenario and value-language resources provide grounded situations and mappings to value categories, but do not make their outputs interchangeable.ValuePortrait connects situated responses to human psychometric profiles.
- Measurement sensitivity: Prior work documents measurement sensitivity to prompts, culture, roles, response format, option order, and evaluator choice.Pairwise choice introduces order effects, while open text introduces evaluator dependence.
- Attitudes and action: Research contrasting stated attitudes with decisions motivates STONIC’s same-item distinction between isolated endorsement and counterbalanced choice.L1 records endorsement; L2 tests whether that ordering predicts a choice.
- Reporting and auditing: Reporting frameworks emphasize conditions and limitations rather than reducing multidimensional systems to one underspecified score.STONIC turns that principle into a concrete decision rule for value profiling.
3 The STONIC Measurement Contract
STONIC uses a common item spine to compare distinct elicitation layers without pooling their outputs. It reports only relations that satisfy explicit coverage, stability, and cross-bank criteria.
- Four Interfaces on a Common Item Spine: L1 independently rates alternatives, L2 presents every defined pair in both orders, and L3 requests a concise free response without alternatives.L2∗ later pairs the model’s strict L3 answer with each authored alternative in both orders.
- Four Interfaces on a Common Item Spine: L0 is a separate 40-item PVQ anchor, while value-discernment tasks are treated as a precondition check rather than another elicitation layer.The anchor portraits do not correspond to scenario items.
- Bank structure: The four scenario banks remain distinct datasets under common prompts and contracts, without treating their source annotations as interchangeable Schwartz labels.Their terrain and source materials differ deliberately.
- Behavioral claims: C1 assigns 1 when a later choice selects the higher-rated alternative, 0 for the lower-rated alternative, and 0.5 for tied ratings.Bank effects receive equal weighting.
- Behavioral claims: C5 scores stable own-answer choices as +1, stable authored choices as −1, and order disagreement as 0, retaining zeros in the denominator.Unparsed comparisons remain missing, while E2 measures P(choose A) −0.5.
- Semantic transfer: C2 compares same-situation ten-value profiles against within-bank item-shuffled profiles across rating–choice, rating–free-answer, and choice–free-answer pairs.The primary semantic view uses signed vectors and alternative scorer views audit dependence on parsing and model family.
- Reporting criteria: Claims require numerical identifiability, at least 80% valid coverage, 80 valid item clusters per bank, and consistent corrected cross-bank direction.Unsupported stronger comparisons remain unmeasured rather than being assigned zero effects.
- Hidden-State Audit: Hidden-state probing stores prompt-end and generated-token positions from one forward pass, using locked test metrics for ratings, coordinates, and choices.Prompt-boundary results support pre-decision interpretation, whereas post-output positions represent the completed response.
4 Experimental Design
The experiment evaluates 35 fixed configurations under preregistered behavioral, semantic, and hidden-state procedures. It preserves missingness and documents the reproducibility boundary of the released analyses.
- Models: The panel contains 35 fixed configurations spanning 22 named checkpoints and 13 matched base–instruction families.Claims are made at configuration level rather than vendor level.
- Models: Base/raw cells remain in the panel because invalid output contracts must be distinguished from absence of value evidence.Base and instruction cells are never compared through imputed values.
- Inference: Inference uses within-bank permutations, clustered behavioral tests, bootstrap confidence intervals, Holm correction for predefined families, and Benjamini–Hochberg correction for exploratory matrices.Banks receive equal macro weight, and parse failures remain missing.
- Reproducibility: The split and inferential registry were generated without reading response or scorer outcomes, and three seed checks changed neither corrected decisions nor confidence-interval inclusion.The prespecified seed remains primary.
- Scorer audit: The scorer audit compares signed primary outputs with DeBERTa classifier views, direct DeBERTa views, and FULCRA regression, while preserving presence and direction separately.A balanced sample of 200 L3 responses receives three Schwartz-10 annotations per response.
- Results reporting: Table 1 counts a configuration as supported only when it is estimable and meets coverage, multiplicity, direction, and cross-bank requirements.R, C, and F denote independent rating, conflict choice, and free answer.
- Reproducibility boundary: The reproducibility boundary excludes model weights, caches, complete raw generations, complete hidden tensors, operational logs, and credentials from the arXiv source bundle.Stored artifacts include prompts, output tokens, parser results, masks, decoding metadata, and hidden summaries.
5 Results
Behavioral continuity is supported for endorsement-to-choice relations, while semantic transfer from ratings to free answers remains numerically positive but insufficiently covered for a stronger shared-profile claim.
- 10 configurations preserve independent endorsement–conflict-choice relations after correction and every cross-bank check.
- Preference transfers more reliably than the inferred value profile across interfaces.
A Profile shape across interfaces B Coherence versus own-answer choice
Profile similarity is strongest between ratings and conflict choices, whereas own-answer preference is widespread despite profile drift; scorer and representation choices affect interpretation.
- A Profile shape across interfaces: 0.73 is the median rating–choice rank correlation, exceeding 0.43 for rating–free answer and 0.50 for choice–free answer.
- B Coherence versus own-answer choice: 0.790 is the median own-answer preference effect among 17 eligible configurations, with effects ranging from 0.508 to 0.932.
- B Coherence versus own-answer choice: 0.0218 is the median signed position-sensitivity effect across 18 qualifying cells, with a range from −0.1136 to 0.1572.
- A Profile shape across interfaces: L1 and L2 profiles are more similar to each other than either is to free text, while L0–L3 is nearly absent at ρ = .037.
- A Profile shape across interfaces: The five semantic views preserve aggregate effect signs in 94.4–100% of comparable profiles but differ in scale, activation density, and individual labels.
- A Profile shape across interfaces: FULCRA most closely matches task-local human-majority labels, while GPV→DeBERTa19 retains ranking information after calibration.
- A Profile shape across interfaces: Post-output hidden states improve test scores over prompt boundaries, indicating decodability of the realized answer rather than a causal value mechanism.
- B Coherence versus own-answer choice: 17 configurations have complete inputs for the exploratory four-bank comparison, but none meets the confirmatory all-bank semantic-coverage requirement.
6 Implications for Alignment Evaluation
STONIC argues that alignment evaluations should preserve distinctions among behavioral and semantic evidence, report their conditions, and avoid unsupported value-identity claims.
- Behavioral and semantic consistency answer separate questions, with behavioral measures providing the clearest cross-interface evidence.
- L3 interpretation depends on the response, parser, scorer, segmentation policy, and missingness rule.
- The exploratory composite ranks 17 configurations, but none passes the confirmatory all-bank semantic-transfer gate.
- STONIC reports eligible effects, coverage, bank stability, order sensitivity, scorer identity, and the corresponding claim in a measurement card.
- Preserving item correspondence and naming the scorer makes changes in evidence visible across policy, decisions, explanations, and representations.
7 Limitations and Ethical Considerations
STONIC limits its conclusions by dataset coverage, parsing, ontology, semantic-audit scope, and the noncausal interpretation of hidden probes. It also treats value profiling as ethically risky when used to label systems or people.
- The four banks broaden domains but do not represent all cultures, languages, or deployment settings.
- Strict parsing creates missing-not-at-random coverage, especially for base models, limiting comparisons.
- Schwartz-10 is an interpretable reporting space, not an exhaustive ontology.
- The semantic audit covers 200 instruction-model responses and three scorer pipelines, calibrating the audit without establishing universal ground truth.
- Hidden probes establish predictive information, not causality or a localized value circuit.
- Value profiling can be misused to label systems, developers, or users as morally desirable or undesirable.
8 Conclusion
STONIC tests whether ratings, choices, and generated language support the same value claim across interfaces. The evidence supports interface-specific behavioral continuity, but not one cross-interface semantic identity for a model.
- Across 35 configurations and four banks, isolated endorsement predicts conflict choice for ten configurations.
- Every eligible model prefers its own generated answer, while option order changes choices.
- Semantic transfer lacks all-bank coverage, and value labels change with the scorer.
- The data support interface-specific behavioral claims rather than one cross-interface semantic identity for a model.
A How to Read the Results
STONIC separates estimability, reporting eligibility, and inferential support, then interprets results within a fixed multi-interface design. Its behavioral effects are positive and reproducible, while semantic conclusions remain constrained by coverage and scorer dependence.
- How to Read the Results: An effect is estimable when structurally valid responses suffice for a numerical effect, meets reporting criteria when prespecified checks pass, and is supported when inferential testing also passes.
- How to Read the Results: 10 of 35 configurations, or 10 of 17 with estimable effects, show independent endorsement predicting later conflict choice.
- How to Read the Results: All 17 configurations meeting reporting criteria support own-answer preference, while all 18 meeting corresponding criteria show presentation-order sensitivity.
- How to Read the Results: None of the 17–23 numerically estimable semantic-transfer configurations meets the prespecified all-four-bank coverage requirement.
- How to Read the Results: L2* pairs each valid L3 answer with every authored alternative in both orders, retaining only pairs whose orders both parse.
- How to Read the Results: C5 ranges from .5076 to .9322 and is positive for all reporting-eligible effects; E2 ranges from −.1136 to +.1572.
- How to Read the Results: The worked GLM-4.7 case demonstrates that rating, conflict choice, and free response are different task observations answering different questions.
I Hidden-State Results
STONIC finds that hidden states encode completed decisions more clearly after generation, while cross-bank profile transfer and semantic agreement remain interface- and scorer-dependent. The resulting terrains and audits support behavioral continuity without a single scorer-independent value profile.
- Hidden-state encoding: Post-output hidden-state features outperform prompt-boundary features in nearly all paired cells, with the largest advantage for L2∗.The L2∗ answer directly reveals the comparison result; only prompt_end supports a pre-decision interpretation.
- Cross-bank transfer: L1 coordinates transfer strongly across banks, L3 moderately, and L2 below its permutation baseline.Behavioral consistency can coexist with weak transfer of a linear boundary across bank-specific coordinate systems.
- Interface profiles: L1 and L2 profiles are most similar, whereas L0 and L3 are almost unrelated because they differ in item identity and elicitation format.The signed L3 mean also separates detection frequency from conditional valence, which a single mean can conceal.
- Interface profiles: Leading values vary by bank, making scenario distribution part of the profile’s scope rather than noise to remove.ValuePortrait emphasizes Benevolence and Achievement; AIRiskDilemmas and MoralChoice emphasize Security; DailyDilemmas emphasizes Conformity and Benevolence.
- Terrain construction: Instruction-model terrains use fixed ten-dimensional Schwartz projections with equal bank weighting and no pooling across models.Complete four-bank terrains are identifiable for 17 models under independent endorsement, 16 under conflict choice, and all 20 under free response.
- Scorer audit: Three-way annotation yields moderate agreement, supporting majority presence as a task-local reference rather than universal ground truth.Binary presence has Fleiss κ = .415 and raw agreement .738; four-state labels have κ = .403 and raw agreement .718.
- Scorer audit: FULCRA most closely matches majority presence, while parser-plus-DeBERTa retains useful rank information after calibration.The held-out-bank F1 for GPV→DeBERTa19 is .629; ValueLlama’s lower presence agreement reflects a different signed relevance-and-direction contract.
- Scorer audit: Direct sentence and whole-text DeBERTa agree strongly, but ValueLlama and FULCRA show substantially weaker row-level agreement with direct sentence scoring.Model-level aggregation increases correlations, so aggregate agreement does not establish per-answer construct validity.
L Exploratory Ranking and Missing Inputs
STONIC’s exploratory ranking combines behavioral and semantic factors only after coverage adjustment, but missing inputs and failed coverage criteria prevent deployment-level interpretation. The ranking therefore remains diagnostic rather than prescriptive.
- Exploratory ranking: Coverage-adjusted cosine is Cd,b = cd,b(1 + rd,b)/2, and the semantic factor averages 12 direction–bank cells equally.The displayed exploratory score combines this semantic factor with the behavioral factor.
- Exploratory ranking: The exploratory score is a post-outcome product of behavioral and semantic factors and lacks an estimable joint bootstrap interval from released aggregates.It does not override the confirmatory reporting criteria.
- Missing inputs: All 17 numerical ranks fail the all-bank C2 coverage requirement, while the remaining 18 cells are structurally non-estimable.Ranks are diagnostic orderings of measured components, not deployment recommendations or moral quality scores.
- Exploratory ranking: Scorer-specific rank correlations with the signed primary view range from .966 to .988, but shared behavioral and coverage factors explain much of this stability.Row-level semantic agreement is substantially weaker.
- Reproducibility: The experiment is deterministic from fixed aggregate artifacts, with explicitly seeded permutation, bootstrap, and generation procedures.Error bars use item-cluster bootstrap intervals.