Source-linked AI summary

Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy

Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos, Anabela C. Areias, Maya D'Eon, Fabíola Costa, Ricardo Rei, Nuno M. Guerreiro

arXiv:2608.21325v1cs.CL

TL;DR

LLMs are increasingly used for emotional support, but their therapeutic conduct remains poorly characterized. The paper introduces and validates a ten-move ontology, applies it to human and model-led sessions, and finds systematic differences that tool-based steering partially narrows. Tools roughly halve mean distributional deviation and improve turn-level alignment by 7–9 percentage points without fine-tuning.

  • Problem

    Research has limited evidence about which therapeutic interventions LLMs favor or neglect and how their psychotherapy conduct compares with trained clinicians.

  • Method

    The paper develops ten function-based moves grounded in MULTI-60, validates them with five licensed psychologists, and applies them to human and synthetic counseling transcripts.

  • Results

    Models over-inquire, neglect psychoeducation, and remain context-anchored, while exposing moves as tools roughly halves mean deviation and improves turn-level alignment by 7–9 percentage points.

  • Takeaways & Limitations

    The ontology provides a shared, auditable vocabulary for specifying and evaluating therapeutic behavior, while tool framing steers models closer to human move patterns without fine-tuning.

  • Takeaways & Limitations

    The study compares spoken human transcripts containing non-verbal information with text-only LLM interactions, and its ontology validation achieved only moderate inter-annotator agreement.

Abstract

from arXiv · show

Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.

1 Introduction

The paper asks how LLMs conduct psychotherapy and introduces a validated ontology for measuring and steering therapeutic behavior. It compares model and human move patterns and finds systematic differences that tools can partially reduce.

  • Users increasingly seek LLMs for emotional support as mental-health needs and service gaps grow.
  • Outcome-level evaluations leave unresolved which therapeutic interventions LLMs favor, neglect, and use relative to trained clinicians.
  • The study introduces ten function-based therapeutic moves grounded in MULTI-60 and validated by five licensed psychologists.
  • The experiments compare human and synthetic counseling transcripts using aggregate move distributions and their temporal structure.
  • Models over-inquire at up to three times the human rate, neglect some moves, and remain strongly anchored to strategies initiated by human clinicians.
  • Exposing the ontology as tools roughly halves mean deviation from human move distributions and improves turn-level alignment by 7–9 percentage points.

2 Background

Psychotherapy process research treats clinician language as observable intervention choices, but existing frameworks differ in scope and modality coverage. MULTI provides a cross-theoretical vocabulary that this paper compresses for turn-level LLM analysis and tool use.

  • Psychotherapy coding frameworks map therapist utterances to discrete strategies such as questions, reflections, interpretations, and challenges.
  • MULTI-60 describes therapist behavior through 60 jargon-reduced items spanning eight therapeutic orientations.
  • MULTI is descriptive rather than a measure of therapeutic adherence or competence, and its session-oriented inventory is difficult to apply densely at turn level.
  • The paper further filters and groups MULTI items into a compact ontology suited to LLM annotation and classification.
  • Each therapeutic move is exposed as a tool whose description specifies clinical appropriateness and whose return value specifies response generation.

3 Ontology

The ontology maps therapist turns to function-based, MULTI-60-traceable moves with multi-label and contrastive coding rules. Five licensed psychologists validate its operational use, while an LLM judge scales classification with moderate expert agreement.

  • Each therapist turn maps to one or more of ten function-based moves grounded in MULTI-60 items for clinical traceability.
  • Coding follows therapeutic function rather than syntax, permits multiple labels per turn, and defines boundaries between commonly confused moves.
  • Five licensed psychologists independently labeled therapist turns across human and LLM-generated corpora using the multi-label ontology.
  • Mean pairwise inter-annotator agreement on human transcripts was 0.627, with pairwise scores ranging from 0.593 to 0.653.
  • A GLM 5.2 judge classifies each turn five times by majority vote and is evaluated against human annotator agreement.

4 Methodology

The methodology compares clinician and LLM therapeutic-move structure across human-grounded and fully synthetic sessions, with and without direct access to the ontology as tools. It also examines how prior transcript context shapes model behavior and whether ontology tools steer models toward more human-like distributions.

  • The analysis measures move statistics over transcript corpora as a proxy for similarity to human counseling practice.
  • The research questions ask how clinician and LLM move structures compare, whether ontology exposure produces more human-like distributions, and whether models remain similar to humans throughout sessions.
  • The experiments also test whether conclusions hold in fully synthetic sessions without human grounding in therapy structure.These sessions remove the human transcript context used in the contrasting setup.
  • The study compares No-Moves prompting with With-Moves generation, where the ontology is introduced as a set of clinical tools.No-Moves relies on a clinician persona system prompt, whereas With-Moves directly grounds generation in the ontology.
  • The model-based clinicians comprise GLM 5.2, Claude Sonnet 4.6, and GPT 5.6 Terra.
  • The analysis models ontology entries as states in a stochastic process, comparing move distributions across turns conditional on prior states and patient utterances.

5 Results

Across human-anchored and model-led counseling transcripts, LLM therapeutic move distributions differ from human clinicians in systematic, context-sensitive ways. The moves framework improves alignment and narrows distributional gaps, but does not eliminate them.

  • 5.1 General Analysis of Move Distribution: Models repeatedly favor Inquiry, while the moves framework reduces this behavior and shifts probability toward Shared Understanding.The largest reduction exceeds 30 p.p. for GPT 5.6 Terra.
  • 5.1 General Analysis of Move Distribution: The moves framework improves parity for Inquiry and Shared Understanding but leaves heterogeneous differences across less frequent moves.Psychoeducation remains consistently neglected across models, including when move tools are available.
  • 5.1 General Analysis of Move Distribution: GLM is consistently closer to humans than the other models, both with and without the ontology.Model-specific profiles differ: Sonnet remains prone to repetitive Inquiry, while Terra is especially susceptible to tool influence.
  • 5.1.1 Alignment Beyond a Single Generation: 62% pass@4 is achieved with the moves framework versus 53% without it under the majority reference.Under the union reference, the corresponding rates are 72% and 65%; the framework advantage persists across k and both reference definitions.
  • 5.1.1 Alignment Beyond a Single Generation: Human and model move trajectories overlap for many moves, but Inquiry, Shared Understanding, and Psychoeducation remain notable exceptions.Temporal comparisons are confounded because some moves are more likely at particular stages of a session.
  • 5.2 Synthetic Transcripts: Removing the human prefix disperses roughly half of Inquiry’s probability mass and increases Interpret, Support Change, and Psychoeducation.Interpret becomes the third most used move, behind Shared Understanding and Inquiry, at more than double the prefixed-model rate and around three times the human rate.
  • 5.2 Synthetic Transcripts: Models rarely initiate Skill Building, with the No-Moves condition producing it in less than <1% of turns.Models often carry forward exercises initiated by human clinicians but rarely start them independently.
  • 5.2 Synthetic Transcripts: Context is a primary driver of LLM therapist behavior: models mirror human strategies when anchored to human turns but adopt a distinct profile when leading sessions.Model-led sessions are interpretation-heavy, quick to propose action, and reluctant to initiate skill work; the moves framework moderates but does not remove this shift.

6 Related Work

Prior work studies LLMs for emotional support and therapy, including response quality, symptom outcomes, and automated coding, but this paper focuses on characterizing therapist behavior through therapeutic strategies.

  • Strategy-guided dialogue research separates sequential strategy selection from response generation, using prompts or external policy planners.
  • Helping-skill frameworks have also been used to create supervised fine-tuning data for emotional-support systems.
  • Controlled trials report symptom reduction, while LLM responses are often rated comparable to or better than human-written responses for emotional support.
  • Automated classifiers have been trained to code therapy transcripts, and related work has classified therapist utterances from GPT-4 and Llama-2 systems.

7 Conclusion

The paper introduces and validates a ten-move ontology for comparing LLM and human psychotherapy behavior, finding systematic differences and measurable improvement when moves are exposed as tools.

  • The ontology contains ten therapeutic moves grounded in MULTI-60 and validated with five licensed psychologists for characterizing LLM psychotherapy relative to human clinicians.
  • Models over-inquire, neglect psychoeducation, and sustain human-initiated strategies while rarely initiating them when leading sessions.
  • Framing moves as tools halves the mean deviation from the human move distribution and raises turn-level alignment by 7–9 percentage points without fine-tuning.
  • The ontology provides clinicians and developers with a shared, auditable vocabulary for specifying and evaluating therapeutic behavior in LLMs.

Limitations

The study’s comparisons are constrained by differences between human and model environments, simulated patients, and only moderate human validation agreement.

  • Human reference sessions are transcribed spoken therapy, whereas the LLMs operate in text, creating a modality mismatch involving access to non-verbal signals and conversational pacing.
  • The LLM-led experiment uses an LLM as the patient, which isolates the therapist from a human context prefix but cannot fully replicate a real person seeking therapy.
  • Ontology validation achieved only moderate inter-annotator agreement, reflecting subjectivity in coding psychotherapeutic dialogue and adding noise to the analysis.

Ethical Considerations

The study does not advocate deploying LLMs as replacements for licensed clinicians. Its ontology is descriptive, and crisis management and safety-critical behavior are outside the analysis.

  • The authors do not advocate deploying LLMs as replacements for licensed clinicians.
  • Closer alignment with human move distributions does not certify clinical competence or safety.
  • Crisis management and safety-critical behavior are outside the scope of the analysis.

A Annotation Procedure

The ontology was developed by psychologists and evaluated through independent multi-label annotation by five licensed clinical psychologists across human and synthetic corpora. Long human transcripts were capped at 60 therapist turns to broaden coverage while controlling annotation costs.

  • The ontology was developed by a team of PhD-level psychologists before the annotation campaign.
  • The annotated data included recorded therapy sessions and 18 synthetic conversations generated with Claude Sonnet 4.6.
  • Human transcripts were capped at 60 annotated therapist turns per transcript to include more transcripts within the annotation budget.
  • Five licensed clinical psychologists independently labeled every therapist turn across three corpora using the ontology.All five annotators coded every turn, and the annotation allowed multiple moves when a turn performed multiple therapeutic functions.
  • Annotators completed feedback-based training on two transcripts excluded from agreement analysis and the main experiments.

A.1 Data Filtering

The study filters human therapy data from the Alexander Street collection and uses synthetic sessions to remove conditioning on preceding human clinician turns. Annotation agreement and move-distribution analyses compare human and model-led transcripts.

  • Human transcripts were drawn from Alexander Street’s Counseling and Psychotherapy collection, retaining CBT and closely related therapy modalities.
  • Table 4 reports pairwise Krippendorff’s α under Jaccard distance for human and synthetic corpora, with all five annotators coding every turn.
  • Fully synthetic transcripts remove the confound of models responding to preceding human clinician moves.Both therapist and patient are simulated, so each clinician turn depends on the model’s prior turns and a simulated member.
  • Figure 6 compares mean move shares across human-led and model-led transcripts averaged over Claude Sonnet 4.6, GLM 5.2, and GPT 5.6 Terra.
  • Synthetic sessions used one simulated member per profile for each clinician model and condition.

D Compute Usage

The experiments used specified computing systems and a tool-based therapeutic-move framework. The appendix documents the ontology’s coding rules, prompts, tool requirements, session constraints, and synthetic-profile distribution.

  • D Compute Usage: Experiments ran on Blackwell Ultra B300 systems or through Vertex AI model-inference credits.
  • The coding manual supports both turn-level annotation and runtime move selection for therapeutic LLM assistants.
  • The system prompt requires the clinician agent to call one or more therapeutic-move tools before every response when the framework is active.
  • Synthetic member sessions enforce a minimum of 40 messages before the conversation can end.
  • The synthetic corpus distributes sessions across 40 member profiles and eight conversation topics.
  • The ontology assigns all explicitly present therapeutic moves in a therapist turn, allowing multiple labels for meaningfully blended functions.

F.4 Decision Order (Function-First)

The judge labels therapist turns by therapeutic function, applying an ordered decision process rather than relying on wording or syntax. It uses session history for context while restricting labels to explicitly present moves.

  • The coding process applies rules from top to bottom, adding each move that is clearly present as a distinct function.
  • The judge assigns one to three permitted labels, preferring a single label unless multiple substantive functions are clearly present.
  • The prompt supplies the ontology manual, allowed labels, preceding session turns, and the therapist turn to be labeled.The judge is instructed to use history only for context and not infer unstated intent.
  • The decision rules distinguish inquiry from interpretation, support for change, shared understanding, and challenge according to each response’s therapeutic function.A tentative understanding seeking confirmation may receive both shared-understanding and inquiry labels when information seeking is meaningful.
  • Function takes precedence over tone and syntax, so a question is not automatically labeled as inquiry.Interpretive, motivational, or process-focused questions receive the corresponding move label when that function is primary.

F.5.1 do_process_alignment

The ontology defines therapeutic moves by their functions in the interaction, including process alignment, goal setting, inquiry, shared understanding, and psychoeducation. Each move is bounded by neighboring functions and illustrated through examples or operational guidance.

  • F.5.1 do_process_alignment: Process alignment calibrates pacing, depth, direction, tone, and collaboration, including repair after the therapist likely creates strain.It is process-focused meta-communication rather than exploration of session content.
  • do_goal_setting: Goal setting collaboratively identifies, clarifies, or prioritizes workable targets, especially when multiple issues compete or session time is limited.It differs from action planning, which specifies concrete steps, and from inquiry, which only gathers details.
  • do_inquiry: Inquiry uses focused questions or instructions to deepen understanding of meaning, sequence, context, and concrete details.Its function is understanding, not persuasion, confrontation, process repair, interpretation, or motivation evocation.
  • do_shared_understanding: Shared understanding communicates accurate tracking through reflection, paraphrase, synthesis, validation, normalization, or brief acknowledgments.It commonly follows disclosure and may slow the pace before interpretation, challenge, planning, or change-oriented guidance.
  • F.5.8 do_psychoeducation: Psychoeducation provides tailored information, rationale, or explanatory framing to improve understanding of experiences, symptoms, treatment processes, or recommendations.The intervention can explain treatment rationales, maintaining mechanisms, skill purposes, and links among thoughts, behavior, physiology, and context.
Loading 2608.21325v1…