Source-linked AI summary
StageWell: A Process-Aligned Chinese Corpus for Positive-Psychology Support Dialogue
Yuxiong Wang, Ziwei Lin, Bo Wang, Yu Zhang, Shiguang Ni
TL;DR
Positive psychology dialogue requires empathetic responses and coherent progression through a multi-turn support process, but existing resources provide limited process-oriented and process-localized supervision. StageWell introduces a process-aligned Chinese corpus and HQS protocol, using targeted repairs for preference learning; across four open-source LLMs, the supervision improves process control, response quality, and safety.
Problem
Existing positive-psychology dialogue resources provide limited supervision for multi-turn process progression and rarely localize process-level defects in preference pairs.
Method
StageWell combines an explicit six-stage support process, a corpus with SFT and DPO data, and HQS-guided preference pairs formed by repairing flawed responses under fixed dialogue history and stage constraints.
Results
Across four open-source LLMs, SFT raises averaged S-exact from 0.424 to 0.646 and reduces H-critical from 0.202 to 0.040, while DPO further improves Q Overall from 3.821 to 3.993.
Takeaways & Limitations
The results support modeling supportive dialogue as a structured multi-turn process, with SFT establishing progression and DPO providing finer local corrections.
Takeaways & Limitations
The six-stage process and HQS criteria are tailored to Chinese-language supportive interaction, limiting direct transfer to other languages, cultures, or specialized scenarios.
Abstract
from arXiv · showhide
Positive psychology dialogue aims to support emotional distress and positive resource building, requiring models to produce not only empathetic replies but also coherent progression through a multi-turn support process. Existing resources often reduce supervision to turn-level strategies or holistic preference labels, leaving process position, support function, and local repair targets implicit. We introduce StageWell, a process-aligned Chinese corpus for positive psychology dialogue, together with HQS, a structured protocol for data construction and evaluation. StageWell organizes support into a six-stage support process and uses a multi-agent whole-dialogue rewriting workflow to construct 12,445 SFT instances, 1,849 DPO preference pairs, and a GroundTruth subset of 120 expert-revised dialogues and 977 QA pairs. Guided by HQS, DPO pairs are built as process-localized repairs: flawed model outputs are used as rejected responses, and targeted rewrites under the same context and stage constraint are used as chosen responses. Across four 9B-14B open-source LLMs, this supervision yields robust gains in process control, response quality, and safety. Averaged across models, BERTScore improves by 0.037, Q-Overall increases by 1.32 points, S-exact increases by 0.236, and the H-critical rate decreases by 0.167. These results highlight the value of modeling supportive dialogue as a structured multi-turn support process rather than as single-turn response generation.
1 Introduction
Positive psychology dialogue requires models to combine empathetic support with coherent progression through a multi-turn process. StageWell addresses gaps in process-oriented and localized preference supervision with a six-stage Chinese corpus, HQS, and targeted rewriting.
- Positive psychology dialogue targets everyday distress relief and positive resource building through multi-turn support involving acknowledgment, clarification, resource activation, and action advancement.
- Existing resources provide limited supervision for multi-turn process progression and rarely localize preference reasons or process-level defects.
- StageWell organizes support into a six-stage process and constructs SFT, DPO, and GroundTruth subsets for Chinese positive psychology dialogue.
- A Planner, Writer, Patcher, and Refiner workflow rewrites generic mental-health dialogues into coherent positive-psychology dialogues, producing 12,445 SFT training instances.
- HQS covers safety, response quality, and stage consistency, supporting both targeted preference repair and structured evaluation.
- StageWell contributes 1,849 DPO preference pairs and provides auditable supervision over support progression and local response quality.
2 Related Work
Related work has advanced empathetic and strategy-aware supportive dialogue systems, but preference supervision still rarely identifies violated process-stage constraints or localized repair targets.
- Recent datasets and systems support empathetic, topic-aware, intervention-like, report-based, long-term, and Chinese counseling dialogue responses.
- The lack of process supervision leaves chosen and rejected responses inheriting ambiguity during preference construction.
- Preference learning resources commonly emphasize instruction following, harmlessness, or broad helpfulness rather than psychological-support process structure.
- Strategy-aware emotional-support methods improve alignment or interpretability, yet rarely identify violated stage constraints, localized defects, or appropriate repairs.
3 Methodology
StageWell models Chinese positive-psychology dialogue as a process-conditioned task, organizing support into six ordered stages and aligning generation with stage-specific functions. Its construction combines whole-dialogue multi-agent rewriting, process-localized DPO repairs, and HQS-based safety, quality, and stage evaluation.
- Task Definition and Six-Stage Support Process: The task requires responses to fit both the current concern and the dialogue’s position within the support process.Warmth alone is insufficient when a response advances too early, repeats empathy, or gives ungrounded advice.
- Task Definition and Six-Stage Support Process: StageWell defines six ordered support stages spanning rapport and safety, clarification, strengths and resources, future and hope, goals and strategies, and consolidation.The stages operationalize per-turn decisions about process position, support function, and progression.
- Process-Aligned Data Construction: StageWell rewrites generic mental-health dialogues into process-aligned positive-psychology dialogues through role, thematic, and risk filtering followed by case-card-guided multi-agent rewriting.The construction workflow is used for dataset creation rather than inference.
- Process-Aligned Data Construction: Whole-dialogue rewriting assigns a global stage sequence, turn functions, and factual boundaries before downstream agents realize and refine each turn.This produces coherent ordered support functions rather than isolated turn edits.
- Process-Aligned Data Construction: DPO pairs fix dialogue history and target stage while contrasting an original flawed output with a targeted repair of a localized defect.This tighter comparison primarily isolates one process-relevant correction instead of varying process position, support function, and surface quality simultaneously.
- Process-Aligned Data Construction: HQS separates hard safety and boundary checks, local response-quality assessment, and stage-consistency analysis for both preference construction and evaluation.Its modules cover safety and risk, five response-quality dimensions, and expected stage, progression pace, and transition fit.
A. Resource scale and basic statistics
StageWell reports resource-scale statistics for its process-aligned subsets and uses HQS to connect construction criteria with held-out evaluation. The GroundTruth subset is disjoint from training resources and serves as expert-referenced evaluation data.
- Resource Scale and Basic Statistics: StageWell’s resource statistics are organized by support stages S1–S6 and HQS response-quality dimensions Q1–Q5.Q1–Q5 denote content fit, acknowledgment, feasibility, maturity, and Chinese naturalness.
- Resource Scale and Basic Statistics: HQS links SFT demonstrations, DPO process-localized preference pairs, and GroundTruth expert references to held-out model evaluation.The protocol is used as a shared framework rather than an isolated metric set.
4 Experiments
The experiments evaluate StageWell through held-out expert references, controlled training conditions, process-aware and reference-based metrics, and a small user-facing pilot. Across four open-source backbones, SFT establishes process control while DPO adds localized refinements, with a small non-monotonic safety case and a preliminary uncontrolled pilot.
- Experimental setup: Experiments use a disjoint GroundTruth set of 120 dialogues and 977 expert-revised QA pairs for held-out evaluation.SFT provides staged demonstrations, DPO provides process-localized preference pairs, and GroundTruth provides expert references.
- Experimental setup: DPO pairs keep dialogue history and target stage fixed while repairing localized defects in rejected model outputs.Earlier-stage repairs address acknowledgment or questioning, whereas later-stage repairs address action grounding, premature closure, or unstable boundaries.
- Evaluation: HQS evaluates safety, response quality, and stage consistency alongside reference-based metrics on the same held-out data.H-critical is dialogue-level, Q scores turn-level quality, and S evaluates predicted support stages.
- Main results: S-exact rises from 0.424 to 0.646 with SFT, while H-critical falls from 0.202 to 0.040 across backbones.These averaged results indicate the largest process-control gains come from stage-aware supervised training.
- Main results: DPO further raises Q Overall from 3.821 to 3.993 and S-exact from 0.646 to 0.660, while lowering H-critical from 0.040 to 0.035.The paper characterizes DPO as local refinement over an SFT-established support process rather than uniform improvement on every backbone.
- Main results: Reference-based gains are modest, and GLM-4-9B is a non-monotonic case where H-critical increases from 0.033 to 0.042 after DPO.The increase corresponds to roughly 4 versus 5 critical cases among 120 held-out dialogues.
- Exploratory feasibility and usability pilot: The pilot reports a higher average immediate psychological-state score from 3.24 to 3.48, a relative increase of 7.4%.It includes 28 students and is intended to assess feasibility, usability, and immediate self-reported experience rather than causal effectiveness.
B. Post-interaction Experience
The pilot’s post-interaction experience results are summarized descriptively, with core ratings near or above 4.0 and lower ratings for helpfulness and continued-use willingness. The study is small, uncontrolled, and short-term, so it provides preliminary feasibility and usability evidence rather than an effectiveness estimate.
- Post-interaction experience: The pilot’s small, uncontrolled, short-term self-report design limits interpretation to preliminary feasibility and usability evidence.It is not designed to estimate causal effectiveness.
5 Conclusion
StageWell combines a process-aligned Chinese corpus with HQS for preference construction and evaluation. Across four open-source LLMs, it improves process control, response quality, and safety while emphasizing stage progression and local repair.
- Conclusion: HQS is a structured protocol for preference construction and evaluation in Chinese positive psychology dialogue.It is presented together with StageWell as a reusable process-oriented resource.
- Conclusion: StageWell provides staged whole-dialogue SFT instances, process-localized DPO preference pairs, and a held-out GroundTruth evaluation set.The corpus targets non-clinical support for everyday distress relief and positive resource building.
- Conclusion: Across four open-source LLMs, StageWell improves process control, response quality, and safety.The conclusion highlights process-aware supervision as practically valuable for supportive dialogue alignment.
- Conclusion: Future work will assess the long-term effectiveness of process-aligned dialogue systems in real-world deployments.The current conclusion frames this as a future research direction.
6 Limitations
The work is scoped to Chinese, non-clinical positive psychology dialogue and uses constrained rewriting, expert revision, targeted DPO repairs, structured evaluation, and a small exploratory pilot. These choices prioritize process supervision and control but limit transferability, naturalistic diversity, preference coverage, and evidence about extended real-world effects.
- StageWell’s six-stage process and HQS criteria may not transfer directly to other languages, cultures, or specialized support scenarios.
- Constrained rewriting and expert revision prioritize clear process supervision and factual control over the conversational diversity of fully naturalistic data.
- Targeted DPO repairs under shared context and stage constraints may cover a narrower range of preference variation than open-ended response comparisons.
- The main evaluation relies on held-out expert-revised data and structured automatic judging, complemented by a small exploratory pilot.
- Larger independent interactive studies are still needed to assess user-facing effects under extended real-world use.
Ethics Statement
The study addresses a sensitive support domain involving vulnerability, privacy, and safety-critical interaction. Ethics approval, deidentification, and manual inspection were used, but black-box models still create risks of subtle bias or inaccuracy affecting vulnerable users.
- All research procedures were approved by a formal ethics committee.
- StageWell underwent deidentification and manual inspection to protect participant privacy.
- Training models on StageWell introduces inherent risks because of machine-learning black-box behavior.
- Subtle model biases or inaccuracies may inadvertently affect vulnerable individuals.
A Additional Dataset Construction Details
The appendix details an auditable, constrained pipeline for constructing StageWell’s student-facing support corpus. It combines stage-aware planning, multi-module rewriting and validation, expert revision, and retained intermediate artifacts across the 12,445-instance SFT dataset.
- Dataset scope: StageWell covers ten common support topics and maps source dialogues to plausible school, family, peer, or growth scenarios.
- Stage-aware rewriting: Each assistant turn receives a dialogue tool, required information slot, and support stage, allowing the rewriting skeleton to remain auditable.
- Pipeline modules: The pipeline combines a Case Card, Planner, Writer, JsonFix, Patcher, Refiner, Validator, and Repeat Gate for structured generation, repair, fluency polishing, and acceptance decisions.
- Pipeline constraints: Planner constraints fix topic assignment, translate adult or workplace concerns into plausible student contexts, and enforce monotonic process progression.
- Auditability: The audit trail retains the normalized Case Card, Planner schedule and facts, labeled Writer dialogue, local edits, and final Validator decision.
- Expert revision: GroundTruth experts revise drafts while preserving topic, turn structure, context, and task scope, targeting expression, empathy, appropriateness, and concrete school-scenario suggestions.
- Dataset checks: The final SFT topic distribution sums to 12,445 instances, while structural diagnostics check that rewritten dialogues remain well-formed.
B.1 Full Training Hyperparameters
The appendix specifies reproducible SFT/DPO training and a separated HQS evaluation protocol. Targeted preference repairs constrain changes to one quality dimension, and chosen responses are often no longer than rejected responses, reducing concern that DPO simply rewards verbosity.
- Training configuration: DPO uses a smaller learning rate of 5 × 10^-7 and larger gradient accumulation than SFT to stabilize preference alignment.
- Checkpoint selection: SFT selects the checkpoint with lowest validation loss, whereas DPO jointly monitors training loss and evaluation metrics, prioritizing the highest evaluation metric when they diverge.
- Training configuration: The four base models share identical primary SFT and DPO hyperparameters to ensure comparability across models.
- Evaluation protocol: HQS separates dialogue-level safety screening, stage-conditioned turn-level quality scoring, and turn-level stage recognition; JSON Repair only recovers formatting.
- Evaluation protocol: The Quality Judge uses negative reminders against rewarding long, polished generic, over-read, bookish, or checklist-like replies.
- Preference analysis: Chosen DPO responses are often no longer than rejected responses, suggesting the preference objective is not simply rewarding verbosity.
- Preference construction: Targeted rewrites improve only designated Qk dimensions while preserving dialogue context, stage function, and core semantics.
- Results visualization: Figure 9 visually compares reference-based and HQS metrics across Base, SFT, and DPO conditions.