Source-linked AI summary

aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI

Michael Keeman, Anastasia Keeman

arXiv:2608.24899v1cs.HCcs.AIcs.CY

TL;DR

No single frontier model can be fully trusted to grade AI safety in the psychological register. The paper compares judges and uses per-metric, anchored selection to support a false-positive-leaning screen that catches 92% of crises.

  • Problem

    No single frontier model can be fully trusted to grade AI safety in the psychological register.

  • Method

    The paper uses per-metric, anchored selection to address structured disagreement among frontier judges and construct a corrected safety-grading target.

  • Results

    92% of crises are caught by a false-positive-leaning screen, while the resulting grader is reported as more faithful than any single frontier judge.

  • Takeaways & Limitations

    Crisis detection is the consistent metric and one that can be acted on, whereas trusting the wrong metric remains a warning.

  • Takeaways & Limitations

    The target is based on one psychologist and is directional rather than a validated gate; comparisons against it are true by construction.

Abstract

from arXiv · show

The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aipsy-bench, an open frozen safety instrument, we run a fully-crossed competence study: three frontier models (gpt-5.4-mini, claude-sonnet-4-6, gemini-2.5-flash) serve as both generators and judges of 3,000 mental-health, companion, and coaching messages against a psychologist's ratings. The disagreement is not noise: it is structured, concentrated on the safety-critical metrics, and one judge (Gemini) is an outlier -- the most lenient, carrying a +0.99 self-preference premium, flagging far fewer tail failures, and scoring a means-in-hand self-harm response "exemplary." Inter-judge agreement on empathy, where sycophancy hides, is the lowest in the battery (alpha 0.24). One axis stands apart: the binary crisis-detection flag is the one safety-critical signal judges agree on (alpha 0.80), erring toward over-flagging, the safe direction for a triage screen. Equal-weight averaging, the canonical fix, blends that leniency and tail-blindness into the safety score. Off-the-shelf open-weight judges are worse for a dispositional, not capability, reason -- and disposition is fine-tunable. We therefore distill a per-metric, psychologist-corrected target into a small, frozen, local model, aipsy-judge-1.0, an Apache-2.0 fine-tune of Gemma-4-26B-A4B. aipsy-judge-1.0 tracks the corrected target better than its base on the composite (ICC 0.64 to 0.75) and crisis detection (kappa 0.65 to 0.82), catches 92% of crises with a false-positive lean, and grades more faithfully than any single frontier judge, while every transcript stays on the machine. These are directional readings against a single-expert-informed target, not validated multi-rater agreement. A safety grader that shares a vendor's post-training shares its blind spots.

1 Introduction

The paper examines whether frontier models can safely judge psychological safety, finding structured disagreement concentrated on clinically consequential metrics. It proposes a psychologist-corrected, local judge while explicitly limiting its claims to a single-expert-informed directional target.

  • The gap: No single frontier model is reliable for grading psychological safety in mental-health conversations.The paper frames this as a clinical-safety gap distinct from general capability and preference evaluation.
  • Frontier-judge findings: +0.99 is Gemini’s self-preference premium on its own generations.The paper also identifies Gemini as an outlier that flags fewer tail failures and rates a means-in-hand self-harm response exemplary.
  • The gap: α = 0.24 is the lowest inter-judge agreement, occurring on empathy where sycophancy can hide.The benchmark treats empathy as clinically contested rather than merely stylistic.
  • Proposed remedy: The proposed remedy selects references per metric against a psychologist’s ratings and distills the corrected target into an open model.The method uses LoRA supervised fine-tuning with stratified failure-oversampling.
  • Deployment: aipsy-judge-1.0 runs locally, keeping sensitive transcripts within the environment and making the grader answerable to a psychologist rather than a vendor.The locality claim is presented as relevant to regulated and PHI settings.
  • Epistemic status: The reported correction is directional because its target comes from one psychologist, not validated multi-rater ground truth.The paper defers numerical validation against multi-rater human agreement to a forthcoming study.

2 Related work and positioning

The paper extends critiques of LLM-as-judge into clinical safety, where structured metric-level disagreement makes simple judge averaging unsafe. It positions psychologist-corrected, per-metric selection as an alternative to both pick-one and equal-weight pooling.

  • Positioning: Clinical-safety disagreement is structured, concentrated on safety-critical metrics, and dangerous when treated as ordinary evaluator noise.This extends prior reliability critiques into a setting involving self-harm and psychological safety.
  • Panel aggregation: Equal-weight pooling fails when one panel member is systematically blind to safety-tail failures.Pooling can reduce variance while preserving a shared directional blind spot and laundering leniency into the mean.
  • Distillation: Sequence-level distillation transfers a corrected rather than raw target into a local judge using parameter-efficient fine-tuning.The paper distinguishes reshaping the signal from cloning a teacher.
  • Prior clinical work: κ = .75 for the best single judge and κ = .74 for the jury can conceal criterion-level collapse to κ = .34.The comparison motivates reporting reliability at metric resolution rather than relying on pooled agreement.
  • Positioning: +0.99 is the own-family premium revealed by this paper’s fully crossed design, which keeps judges disjoint in some prior work but crosses judge and generator families here.The design directly measures self-preference rather than excluding it.
  • Contribution: The paper’s remedy is per-metric psychologist-corrected selection rather than choosing one judge or averaging several.Its comparison is framed as a principled response to structured outlier behavior.

3 Setup: the substrate and the anchor

The study evaluates three frontier judges on a frozen, fully crossed safety benchmark and compares their judgments with a psychologist’s directional reference. The setup separates judge divergence from claims about model safety or validated human agreement.

  • Rubric: Six behaviourally anchored 0–5 metrics plus a boolean crisis-detection flag define the evaluation instrument.The weighted AI-Trust composite places most weight on crisis handling (.30) and empathy (.20).
  • Comparability: The identical frozen evaluator prompt makes model identity the varying factor between judges.The prompt is hash-verified and uses the same retry budget and output parser.
  • Substrate: 3,000 AI messages are drawn from 20 multi-turn scenarios spanning mental-health, companion, and coaching domains.Scenarios escalate from rapport into crisis, enmeshment, or sensitive advice requests.
  • Judge design: The three named judges are gpt-5.4-mini, claude-sonnet-4-6, and gemini-2.5-flash.All run at temperature 0.3 under the same frozen prompt.
  • Judge design: A fully crossed 3 × 3 grid has each judge score every generator, including its own family.The crossing enables a direct self-preference test.
  • Human anchor: A single psychologist scored 201 turns, with 173 complete main items forming the anchor set for psychologist-facing comparisons.The anchor is an expert reference with documented biases, not ground truth or a two-rater consensus gate.
  • Analysis scope: Inter-judge agreement uses all 3,000 items, whereas psychologist comparisons use 173 main items, so neither constitutes validated accuracy alone.The paper reports scale-appropriate coefficients, including Cohen’s κ for the binary flag and ICC(2,1) for ordinal metrics.

4 Part I—The frontier-judge competence map

The frontier judges show structured, safety-critical disagreement rather than interchangeable competence: Gemini is systematically lenient, while empathy and crisis-response grading expose the largest weaknesses. Crisis detection is the notable exception, with strong agreement and errors generally in the safer over-flagging direction.

  • 4.1 Inter-judge agreement: No single frontier judge is fully aligned with the human anchor, and the equal-weight ensemble averages over structured disagreement where safety matters most.The composite three-judge ensemble has α 0.40, while Gemini is an outlier in pairwise comparisons and the safety-critical axes show concentrated divergence.
  • 4.1 Inter-judge agreement: Empathy has the lowest inter-judge agreement (α 0.24), where sycophantic warmth can be mistaken for genuine attunement.Empathy agreement against the psychologist is also near zero, while warm-but-hollow turns receive high judge scores despite low psychologist ratings.
  • 4.2 Leniency, self-preference, and failure detection: Gemini is the most lenient judge, scoring 4.56 on the composite versus GPT-5’s 4.05 and Claude’s 3.67, with a +0.99 own-family premium.On Gemini’s own generations, Gemini scores 4.14 while the other judges score 3.16; GPT-5 and Claude instead show mild self-criticism.
  • 4.2 Leniency, self-preference, and failure detection: Gemini flags substantially fewer safety failures: 9.7% of crisis turns, 5.4% of boundary turns, and 1.5% of advice-safety turns.Claude flags 30.8% of crisis turns and 15.3% of advice-safety turns, while GPT-5 flags 28.6% of boundary-safety turns; equal weighting inherits Gemini’s blindness.
  • 4.3 Crisis detection is the consistent metric—and the one you can act on: Crisis detection is the strongest shared signal: judges show nominal α 0.80, pairwise κ 0.73–0.90, and only 204 of 3,000 items split.Detection rates still differ, but judges largely agree which turns carry a crisis signal, and their errors tend toward over-flagging—the safer direction for triage.
  • 4.5 The failure modes, named and exemplified: Crisis detection remains actionable locally: aipsy-judge-1.0 achieves 92% recall with the same false-positive lean, while graded axes await validation.The paper treats this as directional against a single-rater anchor rather than validated multi-rater agreement.

5 Part II—Off-the-shelf open judges: a disposition gap, not a capability gap

Off-the-shelf open-weight judges underperform for dispositional reasons: they either collapse on required output format or become ceiling-compressed and blind to failures. Gemma-4-26B is the exception, supporting fine-tuning evaluator disposition rather than building capability from scratch.

  • 5 Part II—Off-the-shelf open judges: a disposition gap, not a capability gap: Six of seven open-weight judges fail through either format collapse or failure-blind ceiling compression.Format-collapse models rarely parse mandated chain-of-thought-then-JSON, while clean-format models score nearly everything 4–5 and catch few safety failures.
  • 5.2 Disposition, not scale: Qwen reads sycophantic warmth as empathy, scoring a dependency-cultivating turn 5.0 when the three frontier judges score its boundary safety 0–1.This reproduces the broader pattern in which warmth can mask unsafe relational behavior.
  • 5.3 One open model breaks the pattern: Gemma-4-26B is the sole open model that combines clean formatting with discrimination, reaching composite α 0.65, ICC 0.69, and Spearman 0.82 against the frontier ensemble.It has failure recall 0.31, flags 16.8% of turns where other clean-format models flag approximately 0%, and uses low scores on boundary violations.
  • 5.4 The refined thesis: The open-versus-frontier judging gap is dispositional rather than capability-based, making evaluator stance fine-tunable.Gemma-4 already uses the low end of the scale, so the proposed pivot is sharpening discrimination rather than building a judge from scratch.
  • 5.4 The refined thesis: Helpfulness and agreeableness can produce ceiling-bunching, and failure recall can collapse when a lenient judge never visits the low end of the scale.The paper characterizes this as a cliff rather than a gradual degradation, because safety failures occupy that low-score region.

6 Part III—Psychologist-corrected distillation: the method

The method builds a psychologist-corrected target through per-metric judge selection, then distills it into a frozen local judge while preserving rare safety failures and correcting training defects.

  • 6.1 Constructing the corrected target: Empathy is weighted 0.25/0.60/0.15 across GPT-5/Claude/Gemini, while crisis handling is weighted 0.45/0.45/0.10.Advice safety uses 0.35/0.50/0.15 and boundary safety 0.45/0.40/0.15.
  • 6.1 Constructing the corrected target: Per-metric trust, rather than global averaging, selects judges against a psychologist anchor to construct the corrected distillation target.The target trusts different judges for empathy, advice, boundary, and crisis-related metrics.
  • 6.1 Constructing the corrected target: Bias correction reduces ensemble empathy bias from +0.21 to +0.01 and crisis-handling MAE from 0.74 to 0.52.The corrected blend becomes the target for distillation.
  • 6.2 Distilling the target: LoRA supervised fine-tuning distills the corrected scores into Gemma-4-26B-A4B, with psychologist-scored items held out entirely.The resulting aipsy-judge-1.0 is frozen, fully local, and uses Q8_0 quantization.
  • 6.3 Failure-aware training: Naive 3× oversampling of all failures scrambled per-axis rank-order, whereas stratified oversampling targets only the rare non-advice/boundary failure tail.The target’s failures are concentrated on advice and boundary, so indiscriminate oversampling creates a skewed training prior.
  • 6.5 The v1 →v2 →v3 progression: The progression fixed silent crisis-field omission, over-correction, and quantization issues before shipping aipsy-judge-1.0.Version 2 raised crisis κ from 0.03 to 0.84, while the shipped version improved advice ICC from 0.135 to 0.33 and boundary ICC from 0.55 to 0.735.

7 Results—the local judge versus the corrected target

aipsy-judge-1.0 preserves the corrected target’s intended advantage over the base model, with strongest performance on crisis detection, although advice and boundary remain weaker and the target is engineered rather than clinical ground truth.

  • 7 Results—the local judge versus the corrected target: aipsy-judge-1.0 tracks the psychologist-corrected target better than the base model, preserving the constructed advantage over individual frontier judges.The comparison to any single frontier judge is by construction because the target was engineered from per-metric selection.
  • 7.2 Fit for purpose as a screen: Crisis-detection recall reaches 92%, firing on 14% of turns versus the target’s 12%, a false-positive lean for screening.On a means-in-hand self-harm case, the model flags crisis and identifies the response as too soft because it lacks a clear urgent command.
  • 7.2 Fit for purpose as a screen: Of 335 missed target-failures, 270 involve boundary, 52 advice, and 34 crisis handling; measured against 652 target-failures, crisis-quality misses are about 5%.The axis counts overlap because some turns fail more than one threshold.
  • 7.3 Honest regressions: Advice ICC falls to 0.33 and boundary ICC to 0.735, both below their base lines, while aggregate failure-recall remains 0.49 versus a 0.70 target.Advice is designated the lowest-confidence output and should be used only to flag for review.
  • 7.4 Structural limits: The local judge’s accuracy ceiling is the reweighted ensemble, not clinical truth, and shared teacher failures cannot be removed by further distillation.Empathy sycophancy is identified as a shared failure requiring a new human anchor and dedicated research.

8 The independence and locality argument

The paper argues that a psychologist-corrected local judge provides independence from vendor-specific blind spots and keeps sensitive transcripts inside the evaluation environment.

  • Independence: A safety grader sharing a vendor’s post-training can inherit that vendor’s blind spots and self-preference, while averaging vendors can retain tail blindness.The paper presents third-party correction against a psychologist as the structurally defensible alternative.
  • Independence: aipsy-judge-1.0 is independent in structure and was developed without a vendor relationship, answering to a psychologist-informed target.The study was self-funded.
  • Locality and privacy: Local execution means sensitive transcripts never leave the environment, eliminating egress to secure, redact, or audit.This removes a leak vector and residual re-identification risk from imperfect de-identification.
  • Locality and reproducibility: Accuracy and privacy are separable benefits: the local judge remains useful for either property independently.The model runs fully locally and reproducibly at a fixed seed and temperature.
  • Locality and reproducibility: Fixed-seed reruns provide reproducibility as an honesty mechanism rather than asserted authority.The frozen model extends the benchmark’s reproducibility discipline to judging.

9 Limitations

The evaluation is directional and structurally bounded: it relies on one psychologist, frozen provider snapshots, scripted stimuli, and an engineered target rather than validated clinical agreement.

  • Human anchor: The anchor is one psychologist with documented professional biases, not a two-rater consensus, and the paper claims no validated human agreement yet.Crisis handling has approximately 11 psychologist-scored turns and advice approximately 24, making those rows especially uncertain.
  • Engineered target: The corrected target is an engineered per-metric reference blend, so beating any single frontier judge is true by construction rather than evidence of clinical truth.The empirical question is whether distillation preserves that constructed advantage in a small local model.
  • Metric-specific limitations: Advice is noisy and conservative at ICC 0.33 and should be used for flag-for-review rather than as a verdict.The released model card designates advice as the lowest-confidence output.
  • Non-stationarity: Frontier judges are frozen snapshots even though providers drift, including a mid-study Gemini preview-to-GA swap.Provider non-stationarity motivates a pinned local judge but limits direct temporal comparability.
  • Deployment scope: The stimuli are pre-scripted, point-in-time scenarios across twenty scenarios and three domains, whereas real deployments are longer, adaptive, and broader.The judge is therefore a false-positive-leaning human-in-the-loop screen, not a machine-only safety gate.

10 Discussion and implications

The discussion argues that structured, safety-critical disagreement makes equal-weight judge ensembles unsafe, and proposes psychologist-anchored per-metric selection distilled into a local judge. The resulting design supports reproducible, PHI-safe screening while remaining directional pending multi-rater validation.

  • Equal-weight ensembles are unsafe when judge disagreement is structured and concentrated on safety-critical metrics.
  • Per-metric, psychologist-anchored selection beats both picking one judge and averaging, and can be distilled into one reproducible model.
  • Agreement scalars require marginal-sensitive companions: tone consistency shows an α 0.81/ICC 0.28 split.
  • Frontier-vendor dependence inherits vendor blind spots and self-preference, making structural independence a design requirement.
  • A local, reproducible, PHI-safe screening judge is deployable where a frontier-API judge is not.
  • A false-positive-leaning screen catches 92% of crises and flags the rest for human review, but remains directional until validation.

11 Conclusion

The conclusion finds no single frontier model fully trustworthy for psychological-safety grading because disagreement is structured and an outlier is blind to dangerous self-harm failures. It presents aipsy-judge-1.0 as a frozen local correction, while scoping its claims to a single-expert-informed target.

  • No single frontier model can be fully trusted to grade AI safety in the psychological register.
  • Frontier-judge disagreement is structured, safety-critical, and dominated by an outlier that graded a means-in-hand self-harm response exemplary.
  • Equal-weight averaging blends the outlier’s leniency into the safety score, so the proposed fix is a per-metric psychologist-corrected target.
  • aipsy-judge-1.0 is a small, frozen, fully-local model that grades more faithfully than any single frontier judge and keeps transcripts on the machine.
  • The accuracy claim is scoped to a single-expert-corrected target, with multi-rater validation still underway.
  • Independence and locality are correctness requirements because vendor-dependent safety evaluation inherits vendor blind spots.

12 Availability and provenance

The paper releases aipsy-judge-1.0 under Apache-2.0 as an offline local model, while withholding fine-tuning artifacts and limiting claims about validation.

  • The model is served as a Q8_0 GGUF via Ollama and is the offline default of the aipsy-bench local lane.
  • The paper is the canonical citation for the model, while checkpoints, LoRA adapters, and training logs remain private.
  • The public release includes the servable model and card but does not promise training or validation scripts.
  • The evaluation substrate is the companion benchmark paper, and multi-rater human validation is intended to upgrade the directional readings.
Loading 2608.24899v1…