Source-linked AI summary
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Haonan Dong, Qiguan Feng, Kehan Jiang, Haoran Ye, Xin Zhang, Guojie Song
TL;DR
Agent values lack dedicated benchmarks despite differing from underlying LLM values and posing dataset-, evaluation-, and system-level challenges. Agent-ValueBench evaluates them across executable tasks and harnesses, revealing a cross-model Value Tide that bends under harness pull and deliberate skill steering.
Problem
Agent values lack dedicated benchmarks, while agent evaluation introduces dataset-, evaluation-, and system-level challenges beyond text-only value evaluation.
Method
Agent-ValueBench uses an automated pipeline with expert refinement to create executable environments, value-conflict tasks, and trajectory-level rubrics.
Results
Agent values exhibit a cross-model Value Tide that bends under harness pull and deliberate steering.
Takeaways & Limitations
The findings signal a shift from model alignment and prompt steering toward harness alignment and skill steering.
Takeaways & Limitations
The underlying causal mechanisms of the Value Tide remain unexplored.
Abstract
from arXiv · showhide
Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly drawn growing research attention, and beneath them lie the values silently steering agent behavior. Existing value benchmarks, however, remain confined to LLMs, leaving agent values largely uncharted. From intuitive, empirical, and theoretical vantage points, we show that an agent's values diverge from those of its underlying LLM, and the agentic modality further introduces dataset-, evaluation-, and system-level challenges absent from text-only protocols. We close this gap with Agent-ValueBench, the first benchmark dedicated to agent values. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that cover 28 value systems and 332 dimensions. Every instance is co-synthesized through our purpose-built end-to-end pipeline and curated per-instance by professional psychologists. Each task ships with two pole-aligned golden trajectories whose checkpoints anchor a trajectory-level rubric-based judge. Benchmarking 14 frontier proprietary and open-weights models across 4 mainstream harnesses, we uncover three concerted findings. Agent values first manifest as a Value Tide of cross-model homogeneity beneath interpretable counter-currents. This tide bends non-additively under harness pull, and yet more decisively under deliberate steering via embedded skills. Together these results signal that the agent-alignment lever is shifting from classical model alignment and prompt steering toward harness alignment and skill steering.
1. Introduction
The introduction argues that agent values can diverge from those of their underlying LLM across intuitive, empirical, and theoretical grounds, while existing benchmarks do not evaluate agent values. It presents Agent-ValueBench as the first dedicated benchmark, addressing the dataset and evaluation challenges of agent-value assessment at scale.
- Agent Values Are Not Identical to LLM Values: Agent values are not identical to LLM values: agents interact with environments, incorporate consequential feedback, and make decisions across long horizons, producing divergent value tendencies despite sharing the same underlying model.The paper supports this claim intuitively, empirically through dual-format comparisons, and theoretically through Theorem 1.1.
- Agent Value Evaluation Is Absent and Non-Trivial: Existing value benchmarks evaluate LLMs rather than agents, whose evaluation requires engineered executable environments and value-laden tasks instead of prompts alone.The introduction identifies dataset-level and evaluation-level challenges beyond standard LLM value evaluation.
- Agent-ValueBench: Agent-ValueBench comprises 394 executable environments across 16 domains and 4,335 value-conflict tasks covering 28 value systems and 332 value dimensions.The benchmark is designed to address the identified evaluation difficulties through a dedicated end-to-end approach.
- Agent-ValueBench: The benchmark targets values under conflict, where psychometric theory suggests that values surface most sharply, and provides an open-source pipeline for synthesizing value-oriented environments and tasks.The pipeline is intended to facilitate broader benchmark development.
- Contributions: The introduction frames the work as opening a new research line through theoretical and empirical problem identification, an automated synthesis pipeline, and a comprehensive agent-value benchmark.These contributions collectively address the gap between LLM value evaluation and autonomous-agent value evaluation.
2. Dataset Construction
Agent-ValueBench constructs executable, cross-domain environments, value-conflict tasks, and trajectory-level rubrics through an automated pipeline capped by expert refinement. Its agent-specific evaluation contrasts two value-aligned golden trajectories using observable checkpoints and behaviorally anchored scoring.
- Overview: The benchmark jointly synthesizes executable environments, value-conflict tasks, and trajectory-level rubrics, with every stage capped by per-instance expert-in-the-loop refinement.The pipeline is designed for agents interacting dynamically with grounded environments rather than text-only inputs and outputs.
- 2.1 Environment Construction: Environment construction distills diverse general and safety benchmarks, synthesizes executable programs, and validates them through test-and-repair loops using NFAcc, PPAcc, and NegPAcc thresholds.The final environment pool is further audited by psychologists for fidelity, authenticity, and self-consistency.
- 2.2 Value-Conflict Task Construction: Each value-conflict task provides two golden trajectories favoring opposing values, with checkpoints serving as observable behavioral differentiators rather than relying on survey or multiple-choice responses.Psychologists vet tasks for authenticity, implicit conflict, consistency, discriminability, and reliable end-to-end execution.
- 2.2 Value-Conflict Task Construction: Tasks span 28 value systems and 332 dimensions, pairing sampled environments with non-overlapping value pairs and realizing each conflict through coarse-to-fine executable task synthesis.Each task develops concrete states, tools, task descriptions, and checkpoint trajectories from an initial value-pair specification.
- 2.3 Trajectory-Level Rubrics: Trajectory-level judging uses a cross-task meta-rubric that scores motivationally relevant Attention, Construal, and Execution through behaviorally anchored, task-specific three-level rubric items aligned with each value pole.Psychologists proofread and rewrite unsound rubric items before deployment.
3. Experiments
Across 14 agents, values converge into a structured Value Tide with shared adherence and priority profiles, interpretable model-specific counter-currents, and selective rather than uniformly prosocial dimensions. Harness substitution bends this tide non-additively in model-dependent ways, while skill steering shifts priorities more deeply than prompt steering and can recover a full target ordering.
- Value Tide: Across 14 agents, adherence converges within narrow system-specific bands and priority ranks collapse toward shared consensus orderings, forming a population-wide Value Tide.Adherence ranges are [5.94, 6.27] on MFT08, [5.16, 6.17] on PVQ40, [5.39, 6.64] on HEXACO, and [5.06, 5.92] on LVI; priority consensus is shown in Fig. 3 and Table 1.
- Value Tide: The shared profile is structurally selective: it elevates safety, autonomy, universalism, and conscientious loyalty while suppressing hedonism, privacy, honesty-humility, authority, and purity.Loyalty crests at 8.01 in MFT08, while PVQ40 highlights Security at 6.93 and Self-Direction at 6.80 and reaches a Privacy floor of 3.59.
- Value Tide: Model-specific counter-currents remain coherent beneath the tide, including Qwen3 30B A3B’s lowest adherence on three systems and localized reversals such as Care versus Purity.These drifts are localized to one or two dimensions and remain traceable to the same agent across adherence and priority axes.
- Harness substitution: Harness substitution changes both adherence and priority substantially: 74% of 93 within-model ranges exceed 1.0, with effects comparable to inter-family differences under vanilla ReAct.The same harness can drive opposite drifts across models, such as Codex moving Closed-Mindedness from 5.96 to 3.29 on GPT-5.4 but from 4.18 to 7.67 on Kimi K2.5.
- Skill steering: Skill steering places an average of 3.7 of 5 values at their exact target rank versus 1.3 under prompt steering, with Codex recovering GPT-5.4’s full target ordering.The recovered ordering is Care≻Authority≻Purity≻Loyalty≻Fairness.
4. Conclusion … C. Notation
Agent-ValueBench is presented as the first benchmark for agent values, revealing a Value Tide shaped by harnesses and deliberate steering while motivating harness alignment and skill steering. The paper also defines ethical safeguards, situates the benchmark within value research, and documents supplementary materials and notation.
- 4. Conclusion: Agent-ValueBench is introduced as the first benchmark dedicated to agent values, with results revealing a Value Tide that bends under harness pull and deliberate steering.These findings signal a shift from model alignment and prompt steering toward harness alignment and skill steering.
- 4. Conclusion: The conclusion identifies unexplored causal mechanisms underlying the Value Tide and calls for future work to investigate them.The benchmark exposes the phenomenon but does not yet explain its mechanistic underpinnings.
- (Appendix): The appendix provides ethics, related-work, notation, implementation, value-inventory, comprehensive-results, prompt-template, proof, and validation-study materials.Listed supplementary sections include Comprehensive Experimental Results, Proof of Agent–LLM Value Non-Identity, and Human Validation, Executability, and Stability Studies.
- A. Ethics Statement: Agent-ValueBench is framed as a descriptive, human-supervised audit rather than a normative authority, universal moral taxonomy, deployment certification, or direct optimization target.Its artifacts include synthesized executable environments, synthetic task states, psychologist-curated conflicts, golden trajectories, and task-specific rubrics, without private logs, surveillance, biometric, or personally identifying behavioral data.
- A. Ethics Statement: The principal ethical risk is dual use: benchmark measurements could be misused to game evaluations or steer agents toward harmful orientations.Recommended safeguards include transparent reporting, independent safety review, sandboxed execution with resource limits, and restricted downstream use.
- B. Related Work: Psychological value theories treat values as cross-situational motivational goals organized into priority structures shaped by compatibility and conflict.Schwartz’s theory emphasizes circumplex relations in which adjacent values are compatible and opposing values express systematic tensions.
- B. Related Work: Existing LLM value evaluations use questionnaires, scales, surveys, and value-conflict trade-offs, while autonomous-agent values remain largely unexplored [19, 20, 91, 92, 93,…].Conflict paradigms examine value prioritization, preference stability, and context sensitivity across dilemmas, role conflicts, social decisions, and high-stakes situations [30] [106] [108] [109] [110] [111] [112] [113].
- C. Notation: Table 2 consolidates notation for environment auditing, task and value indices, conflicting value pairs, executable tasks, and cross-task rubric synthesis.Symbols include NegPAcc, ℓ, eℓ, sℓ, P_s, {v_a, v_b}, cℓ, M, and Φrubric.
D. Implementation Details
The implementation converts trajectory-level adherence into pairwise value preferences, then estimates value priorities with Bradley-Terry models while fixing evaluation and synthesis settings.
- Hyperparameter Settings: Evaluation uses temperature 0.0, with trajectories capped at 50 steps and 12,000 tokens per step, while specified models handle environment synthesis, task synthesis, rubric generation, and judging.GPT-4.1 and GPT-4.1 Mini synthesize environments; Gemini 3.1 Pro Preview synthesizes tasks; DeepSeek V3.2 generates rubrics and serves as the rubric-based judge.
- Value Priority Computation Details: Value priorities are computed by comparing trajectory-level adherence scores for each pair of value dimensions, with ties counted as 0.5 wins for both sides.These comparisons produce effective win counts for every observed value pair.
- Value Priority Computation Details: A separate Bradley-Terry model is fit for each agent–value-system configuration, using positive value-strength parameters to represent relative prioritization.A weak symmetric smoothing term α=0.01 is applied only to existing comparison edges to prevent degenerate estimates from complete separation.
E. Value-System Inventory · BARCHARD
The inventory spans 28 value systems and 332 dimensions, combining broad motivational values with cognitive styles, personality traits, social beliefs, and moral foundations. It includes both established value families and finer-grained dimensions describing how agents interpret, evaluate, and navigate social situations.
- E. Value-System Inventory: The inventory covers 28 value systems and 332 value dimensions across the benchmark’s value space.Its scope is summarized as a complete value-system inventory.
- E. Value-System Inventory: Core motivational values include tradition, benevolence, universalism, self-direction, stimulation, hedonism, achievement, power, and security.The inventory also groups these into higher-order constructs such as self-transcendence, self-enhancement, openness to change, and conservation.
- E. Value-System Inventory: Finer-grained social values cover dependability, caring, tolerance, concern, conformity, security, face, power, and self-direction.These dimensions distinguish interpersonal, societal, personal, resource-based, dominance-related, action-oriented, and thought-oriented forms of value.
- E. Value-System Inventory: Cognitive-style dimensions characterize analytic versus holistic thinking, field versus parts attention, interactionist versus dispositionist causality, and cyclic versus linear change.The inventory also includes contrasting approaches to resolving contradictions, including naive dialecticism and formal logic.
- E. Value-System Inventory: Additional systems capture interpersonal ethics, affect, social worldview, fate, religiosity, and beliefs about competitive versus cooperative exchange.Examples include sincerity, fairness, greed avoidance, modesty, fearfulness, social cynicism, social complexity, fate control, religiosity, zero-sum beliefs, and joint-profit exchange.
- E. Value-System Inventory: Personality-related dimensions span emotionality, extraversion, agreeableness, conscientiousness, openness to experience, and related facets.Examples include fearfulness, anxiety, dependence, sentimentality, social boldness, forgiveness, flexibility, patience, diligence, prudence, inquisitiveness, creativity, and unconventionality.
- E. Value-System Inventory: The inventory incorporates moral foundations and fairness-related intuitions, including care, fairness, loyalty, authority, purity, equality, and proportionality.These dimensions address harm avoidance, equal treatment, merit-based rewards, ingroup cooperation, deference to legitimate authority, and contamination or sanctity concerns.
F. Comprehensive Experimental Results
This section reports the complete experimental results for value adherence and value priority across value systems, models, harnesses, and steering settings.
- F. Comprehensive Experimental Results: The experiments evaluate both value adherence and value priority.
- F. Comprehensive Experimental Results: The results cover multiple value systems and models.
- F. Comprehensive Experimental Results: The evaluation spans multiple harnesses and steering settings.
F.1. Experimental Results on Value Adherence and Value Priority
This section reports Agent-ValueBench’s complete experimental results for value adherence and value priority across all 28 incorporated value systems, presented in Tables 4–22.
- F.1. Experimental Results on Value Adherence and Value Priority: Complete adherence and priority results span all 28 value systems incorporated in Agent-ValueBench.The results are presented across Tables 4–22.
- F.1. Experimental Results on Value Adherence and Value Priority: Tables 4–9 report adherence and priority results for mft08, mpq, bis_bas, csf, ahs, nfcc1993, buss1980, and vsm13.For adherence, bold and underline identify the top-performing model and runner-up per dimension; for priority, they identify first and second ranks within each model.
- F.1. Experimental Results on Value Adherence and Value Priority: Across the tables, bold and underline consistently mark the top two outcomes under the section’s adherence and priority ranking conventions.For adherence, the markers denote the top-performing model and runner-up per dimension; for priority, they denote first and second ranks within each model.
F.2. Additional Harness Experiment Results
This section reports value-adherence and value-priority results for Claude Sonnet 4.6, GPT-5.4, and Kimi K2.5 across four value systems and three agent harnesses. Tables 23–30 present the corresponding model–harness comparisons, including best and runner-up adherence configurations and first- and second-ranked priority dimensions.
- F.2. Additional Harness Experiment Results: The experiments evaluate Claude Sonnet 4.6, GPT-5.4, and Kimi K2.5 in Vanilla ReAct, Claude Code, and Codex across MFT08, NFCC2000, PVQ40, and VSM13.The reported metrics are value adherence and value priority.
- F.2. Additional Harness Experiment Results: For value adherence, bold and underlined entries identify the highest-scoring and runner-up model–harness configurations per dimension.For value priority, bold and underlined entries identify the first- and second-ranked dimensions within each model–harness configuration.
F.3. Complete Prompt/Skill Steering Results
This section reports complete prompt- and skill-steering results for MFT08 and PVQ40, pairing each guided condition with its unguided harness baseline. MFT08 target value orderings are specified for Sonnet 4.6, GPT-5.4, and Kimi K2.5.
- MFT08: The complete MFT08 steering results pair every prompt- or skill-guided condition with its unguided harness baseline in Tables 31 and 32.The reported target orderings are Sonnet 4.6: Purity≻Authority≻Fairness≻Care≻Loyalty; GPT-5.4: Care≻Authority≻Purity≻Loyalty≻Fairness; Kimi K2.5: Authority≻Fairness≻Purity≻Care≻Loyalty.
- PVQ40: The complete PVQ40 steering results likewise pair each prompt- or skill-guided condition with its unguided harness baseline in Tables 33 and 34.The passage introduces target orderings for the evaluated models, but the supplied excerpt provides only the continuation containing those orderings.
F.4. Agent–LLM Priority Comparison · G. Prompt Templates
The paper contrasts value prioritization in identical backbones operating as standalone LLMs versus agents, while documenting prompt templates that construct cases and score trajectories. Its pipeline anchors comparisons in executable task contexts, dual value-aligned trajectories, and auditable rubric-based judging.
- F.4. Agent–LLM Priority Comparison: The LLM-versus-agent comparison includes Gemini 3.1 Pro Preview and GPT-5.4 as standalone LLMs alongside their agent settings.
- G. Prompt Templates: The prompt-template section documents templates for case construction, rubric synthesis, and trajectory-level evaluation.
- F.4. Agent–LLM Priority Comparison: Table 35 compares value-priority rankings between standalone LLMs and agents using identical backbone models across mft08 and nfcc2000.The LLM protocol converts Agent-ValueBench tasks into multiple-choice prompts by combining task descriptions with environmental contexts and offering divergent value-favoring behaviors as options.
- G.1. Task Construction Prompts: Stage 1 drafts realistic operational tasks whose structural constraints induce value conflict without explicitly naming the competing values.The template requires multiple independent decisions, valid tools and dependency-complete state keys, and two feasible but non-identical action paths.
- G.1. Task Construction Prompts: Stage 1 independently specifies checkpoint hypotheses for both value tendencies, including related tools, concrete actions, and observable signals.The template forbids cascading decision lock-in, mirrored A/B templates, and superficial differences in action intent.
- G.1. Task Construction Prompts: Stage 2 realizes each draft into an executable case by co-designing realistic initial state values, tool selections, task descriptions, and operationally distinct value-consistent trajectories.It enforces authoritative schemas and dependencies, justifies intentionally empty states, preserves cross-state consistency, and keeps divergence grounded in state, resource, risk, and timing structure rather than wording.
- G.2. Rubric Synthesis and Online Judging Prompts: Rubric synthesis freezes a case-specific dual-track rubric so an online judge can score one complete agent trajectory against both value tracks consistently and audibly.The judging rules prioritize state changes and tool outputs over tool calls and reasoning text, require rubric anchors to be followed in order, and treat keywords alone as weak evidence.
H. Proof of Agent–LLM Value Non-Identity … I.2. Automated Executability Audit
The paper proves that agent and LLM values coincide only when their induced value-evidence distributions match, while Agent-ValueBench’s retained artifacts show strong validation support and broad automated executability after repair.
- H. Proof of Agent–LLM Value Non-Identity: For a fixed model-side law μ, text-only and agentic value profiles are identical exactly when their measurement channels induce the same distribution over normalized value evidence.If the evidence laws differ, a bounded evidence-scoring rule separates the profiles.
- H. Proof of Agent–LLM Value Non-Identity: Uniform profile equality over every probability law is stronger: it holds exactly when the text-only and agentic evidence kernels coincide pointwise.The proof uses point-mass model-side laws to establish kernel equality and then obtains equality for every input law.
- H. Proof of Agent–LLM Value Non-Identity: Thus, sharing the same model-side source does not guarantee identical values; identity additionally requires equality of the evidence law induced by the two channels.Agent values are therefore not reducible to raw LLM text values by default.
- I. Human Validation, Executability, and Stability Studies: Agent-ValueBench operationalizes value measurement through executable environments, implicit conflicts, A/B golden trajectories, checkpoint anchors, task-specific rubrics, and scalable LLM-as-Judge scoring.The validation argument follows a Messick-style framework.
- I.1. Validation Overview and Population-Level Evidence: The validation package combines complementary checks of engineering validity, artifact plausibility and value discrimination, golden-trajectory semantics, rubric traceability, and judge agreement.The supplied overview identifies seven complementary checks across the validation package.
- I.1. Validation Overview and Population-Level Evidence: More than 94% of retained environments and tasks met the pre-specified expert-quality threshold, while blinded studies supported golden-trajectory polarity, rubric validity, and judge alignment with human application.Validation covered design quality, semantic and operational golden-trajectory roles, rubric traceability, and scoring-layer content validity.
- I.2. Automated Executability Audit: 4,335 retained tasks across all 394 environments were smoke-tested; 4,333 reached terminal states, 4,331 satisfied finite-state completion, and no retained benchmark-attributable failure remained after repair.The formal audit covered 60,480 evaluation rollouts, with deterministic retry reducing execution-attributable invalidity, although the supplied passage does not provide the final reduction value.
I.3. Environment and Task Design Reliability · I.4. Golden-Trajectory Value Anchors · I.4.1. Value-Item Instantiation and Blind Discriminability
Expert review found the retained environments and tasks plausible, coherent, and value-discriminative, while blind evaluation confirmed that golden trajectories instantiate intended values with observable behavioral evidence and minimal confounding.
- I.3. Environment and Task Design Reliability: Environment ratings were high for fidelity, realism, and tool/resource consistency, supporting plausible and coherent interaction substrates.Experts reviewed source-environment summaries and synthesized specifications using structured ratings of structural preservation, realism, and mutual tool/resource coherence.
- I.3. Environment and Task Design Reliability: Task ratings supported realism, implicit conflict design, value consistency, and behavioral discriminability across the retained task set.Reviewers examined task descriptions, environment summaries, available functions, value definitions, and A/B checkpoint anchors.
- I.3. Environment and Task Design Reliability: Task-design executability was interpreted at the specification level, whereas code-level execution was established separately through the automated audit in Section I.2.This distinction prevents expert ratings of specification coherence from being treated as direct evidence of runtime execution.
- I.4. Golden-Trajectory Value Anchors: The blind semantic protocol tested whether two golden trajectories expressed their intended values without cross-value leakage, while rollout coverage tested whether anchors remained active in real non-failed rollouts.Every retained task and golden trajectory was reviewed during construction, with the confirmatory blind semantic study estimating reliability.
- I.4.1. Value-Item Instantiation and Blind Discriminability: Six psychologists evaluated 120 retained tasks and both golden trajectories per task under blinded conditions, excluding a three-rater pilot from headline estimates.Intended labels, checkpoint labels, rubric items, judge scores, model identity, and generation metadata were hidden.
- I.4.1. Value-Item Instantiation and Blind Discriminability: Across 1,440 expert trajectory ratings, 97.1% cited behavioral evidence spans, while primary-basis confounding was 2.9% for both value-item instantiation and blind assignment.The acceptance logic also required high intended-value ratings, low opposing-value ratings, large discriminant margins, reliable panel estimates, and high blind A/B recoverability.
I.4.2. Anchor Coverage in Real Rollouts · I.5. Rubric Content Validity and Traceability · I.6. LLM-as-Judge Validation under Rubrics
Validation studies support the benchmark’s anchor coverage, rubric validity and traceability, and reliability of LLM-as-Judge scoring under frozen rubrics. Real rollouts broadly expose intended anchors, expert audits confirm rubric fidelity, and human–LLM agreement is high.
- I.4.2. Anchor Coverage in Real Rollouts: The audited rollout sample spans all 28 value systems, 16 environment domains, 14 models, and 4 harnesses, supporting broad anchor-coverage assessment.The analysis combines internal-consistency evidence from frozen rubrics synthesized from the anchor system with independent semantic evidence from blind discriminability.
- I.4.2. Anchor Coverage in Real Rollouts: 89.3% ± 8.4% mean checkpoint-opportunity coverage, 96.2% of rollouts covering at least 75% of opportunities, and 92.7% human–anchor agreement validate anchor coverage in real rollouts.Only 2.8% of rollouts fall below 70%, while anchor margin correlates with rubric priority margin (Spearman 𝜌= 0.82, Pearson 𝑟= 0.79, Kendall 𝜏= 0.68).
- I.4.2. Anchor Coverage in Real Rollouts: Human–rubric three-way agreement reaches 91.3% in a 300-rollout holistic audit, complementing the 92.7% human–anchor agreement.These audits provide convergent evidence that observed rollout behavior aligns with both the checkpoint anchors and frozen rubric judgments.
- I.5. Rubric Content Validity and Traceability: The rubric-validity study evaluates 60 stratified rubrics using post-revision Round-2 ratings, covering value systems, value-pair types, environment domains, and rubric lengths.The assessment targets coverage, relevance, clarity, evidence-groundedness, and non-redundancy for formal trajectory-level evaluation.
- I.5. Rubric Content Validity and Traceability: Post-revision rubrics satisfy standard and stricter content-validity criteria, with complete source-checkpoint coverage and high expert semantic-audit rates.The audit reports side fidelity 99.7%, behavioral observability 99.4%, item atomicity 98.4%, non-redundancy 97.3%, and score-anchor clarity 98.1%.
- I.6. LLM-as-Judge Validation under Rubrics: The judge-reliability study uses three blinded psychologists to score 100 stratified task–trajectory pairs with frozen rubrics, comparing agreement and threshold sensitivity.The protocol requires observable evidence, exact rubric application, lower defensible scores under ambiguity, and a pre-specified tolerance 𝜖= 0.50 for priority leaning.
- I.6. LLM-as-Judge Validation under Rubrics: Across 1,080 rubric items, exact human–LLM agreement is 91.8% with mean absolute error 0.09 on the 0–2 item scale.The three-way confusion matrix correctly identifies 41 A-leaning, 12 near-tie/mixed, and 39 B-leaning cases out of 100.
I.7. Computational Stability
Computational stability checks show that aggregate value profiles and directional effects remain stable across rollout stochasticity and judge replay, despite expected nonzero rollout variability. Complementary validation studies assess reliability, trajectory anchors, rubric traceability, and human–LLM judge agreement.
- Stability protocol: Rollout stability repeats stratified task subsets across the main model set, harness comparisons, and steering settings, while judge replay stability reruns the LLM-as-Judge five times on fixed trajectories, rubrics, and prompts.These procedures are reported in Tables 49 and 50.
- Stability results: Aggregate value profiles, priority rankings, harness effects, and steering effects remain directionally stable across repeated rollouts and judge replays, rather than reflecting a single execution or judge call.Rollout variability remains nonzero, as expected for agentic behavior.
- Validation argument: Validation studies use panel-mean artifact scores with ICC(2,k) reliability and assess golden-trajectory ratings, anchor coverage, CVI, checkpoint-to-rubric traceability, and human–LLM rubric-score agreement.Priority is derived from Δ = A−B under the pre-specified near-tie criterion.