Source-linked AI summary
Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts
Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li
TL;DR
Because distress judgments reflect community norms, the paper asks whether LLMs align with the communities whose language they assess. It uses perspectivist Reddit annotations and finds systematic model miscalibration, especially inflated low-severity judgments, with implications for deployment.
Problem
The paper addresses limited evidence on whether distress-detection systems capture pluralistic, community-shaped judgments rather than a single undifferentiated standard.
Method
The study collects community-specific distress and support-seeking aggregates from 321 participants’ judgments on 1,198 Reddit posts, then evaluates LLM configurations against them.
Results
Open-weight LLMs systematically over-estimate distress relative to human aggregates, while GPT-5 and Gemini 2.5 Pro inflate none-to-mild cases and Claude Opus 4 is conservative.
Takeaways & Limitations
Community-specific validation, calibration, and human oversight are warranted because model severity judgments are not uniformly well calibrated across systems or communities.
Takeaways & Limitations
The study measures community perceptions rather than clinical definitions of distress, limiting direct comparability with clinical screening tools.
Abstract
from arXiv · showhide
Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection are typically aligned to a single, undifferentiated standard. How well do these models capture the perspectives of the communities whose language they assess? We address this question through a perspectivist annotation study in which 321 participants provided 9,587 judgments on 1,198 Reddit posts spanning six identity-based communities, yielding community-specific labels. Raters in the contextualized in-group condition show a modest tendency to agree more with their community than uncontextualized out-group raters (OR = 1.18), an effect varying significantly across communities. We then evaluate nine open-weight LLM configurations and four frontier configurations against these labels. Open-weight LLMs systematically over-estimate distress: when communities perceive none-to-mild distress, these models achieve only 31-44% accuracy, predominantly producing false positives. GPT-5 and Gemini 2.5 Pro show the same none-to-mild inflation even when their full-sample over/under rates are mixed, while Claude Opus 4 is more conservative. This pattern does not simply mirror an outsider reading position: uncontextualized out-group human aggregates were nearly symmetric, with 18% over-estimation versus 19% under-estimation. Instead, the models that inflate none-to-mild cases exhibit a distress prior that exceeds both contextualized in-group and uncontextualized out-group human judgments. These findings have implications for equitable AI deployment in mental health contexts, where miscalibrated distress detection may unevenly affect the communities being assessed.
1 Introduction
The paper treats distress and support-seeking judgments as socially situated and evaluates whether community-grounded human labels reveal systematic differences in LLM alignment.
- Motivation: Distress judgments vary with community norms, lived experience, and contextual interpretation rather than a single universal standard.The paper argues that identity-blind pipelines risk missing or misreading cues across communities.
- Research focus: The study examines distress and support-seeking across men, women, non-binary individuals, veterans, mothers, and fathers.These dimensions are related but distinct forms of mental-health expression.
- Research questions: The authors compare contextualized in-group judgments with uncontextualized out-group judgments to establish community-specific alignment targets.The research questions ask how communities differ and whether reading condition affects agreement with community aggregates.
- Contributions: In-group raters were modestly more likely to agree with their community’s distress aggregate than out-group raters, with OR = 1.18 and significant variation across communities.The effect was not reducible to prior subreddit experience.
- Contributions: Open-weight LLMs systematically over-estimated distress relative to both in-group and out-group judgments, while GPT-5 and Gemini 2.5 Pro inflated none-to-mild cases and Claude Opus 4 was conservative.Identity and contextual conditioning reduced over-estimation but could also produce systematic under-estimation.
2 Related Work
Prior work documents variation in how distress is expressed and perceived, but mental-health datasets rarely preserve these perspectives; persona-based alignment offers a possible but fragile response.
- Mental well-being distress detection: Mental-health detection progressed from lexicon-based classifiers toward transformer models and instruction-tuned LLMs.Earlier studies commonly used social-media text and conventional machine-learning models.
- Mental well-being distress detection: Low-to-moderate risk cases are difficult to classify because indirect signals and cultural variation make distress expressions subtle.Recent shared-task analysis linked reduced performance in these cases to indirect linguistic cues.
- Socio-demographic differences: Distress expression differs across demographic groups, including gendered patterns in emotional expression, help-seeking, and diagnosis.The paper motivates studying men, women, and non-binary individuals alongside other identity-based communities.
- Rater identity biases: Emotion-recognition research describes an in-group advantage in identifying emotions expressed by members of one’s own group.Shared identity may help observers interpret subtle cues and community norms.
- Perspectivist annotation: Mental-health distress datasets generally collapse judgments from undifferentiated annotators into identity-blind labels.The paper argues that this can mislabel or overlook signals from particular groups and propagate bias downstream.
- Persona-based design: Persona-based prompting can shift model outputs toward subgroup norms, but its alignment benefits are fragile.The related work presents identity cues as a potential strategy without treating them as a complete solution.
3 Methodology
The study combines purposive Reddit sampling, perspectivist annotation, mixed-effects modeling, and LLM evaluation against community-specific aggregate labels.
- Dataset: Posts came from six identity-oriented subreddits spanning an 11-year period and were filtered for textual content, engagement, and other quality criteria.The communities were treated as approximate contexts rather than exact demographic containers.
- Dataset: Weak supervision used Snorkel heuristic labeling functions to retrieve posts likely to contain distress or support-seeking language.The source communities were not primarily mental-health forums, so retrieval was improved before annotation.
- Annotation Task: U.S.-based adult English speakers were recruited through identity pre-screening, with in-group status defined by demographic match to the post’s source community.Out-group status was defined as not matching the focal community identity.
- Annotation Task: Out-group raters saw post text without source-community information, whereas in-group raters received the subreddit context.The design estimates a contextualized in-group versus uncontextualized out-group reading effect rather than a fully crossed identity-by-disclosure design.
- Annotation Task: 9,587 judgments from 321 participants produced aggregates for 1,198 of 1,200 sampled posts.Raters labeled distress severity, support-seeking when distress was present, and confidence.
- Analytic strategy: Cross-classified mixed-effects logistic regression modeled distress and support-seeking, allowing the in-group contrast to vary by community.Community-specific contrasts were tested against reduced models, with Holm-adjusted p-values for multiple comparisons.
- LLM Alignment: The LLM evaluation used the same 1,198 posts and compared nine open-weight configurations plus four frontier configurations against community aggregates.Models were evaluated primarily with macro-F1, with vanilla and contextualized prompting variants included where applicable.
- Analytic strategy: Dawid–Skene aggregates were derived from at least four in-group or out-group ratings per post, with leave-one-out recomputation for in-group focal raters.The model down-weights inconsistent annotators and estimates latent labels plus rater-specific confusion patterns.
4 Results
Human judgments show a modest, uneven contextualized in-group alignment advantage for distress, while LLM alignment and error patterns vary by model and community. Contextualization improves some models modestly but can shift errors toward under-estimation, and exploratory rationale analyses remain interpretive.
- Human annotation results: κ = .44 for severity and κ = .64 for support-seeking summarize moderate and substantial agreement, respectively, between IG and OG raters.Severity agreement was lowest in r/NonBinary (κ = .39), while support-seeking agreement was highest there (κ = .77).
- Human annotation results: OR = 1.18 indicates that contextualized IG raters had about 18% higher odds than OG raters of matching the distress aggregate.The pooled raw agreement was 82.3% for IG and 80.5% for OG, with significant heterogeneity across communities.
- Human annotation results: OR = 2.96 was observed for IG versus OG distress-aggregate agreement in r/NonBinary in the random stratum.The random stratum strengthened the pooled effect to OR = 1.32 and produced a significant omnibus interaction, whereas the community carrying the benefit was not stable across samples.
- LLM alignment: Qwen3-30B-A3B Thinking achieved the highest open-weight F1 across conditions, while Olmo-3-7B-Base performed worst.Qwen3-30B-A3B Thinking reached C: 0.621 and ∅: 0.612; Olmo-3-7B-Base reached ∅: 0.389.
- LLM alignment: 42% over-estimation versus 4% under-estimation for Olmo-3 and 33% versus 7% for Ministral-3 demonstrate systematic upward distress errors in open-weight models.OG human aggregates were nearly symmetric at 18% over-estimation and 19% under-estimation, indicating an inflated model distress prior beyond outsider reading alone.
- LLM alignment: 40%–42% over-estimation made r/NonBinary, r/TwoXChromosomes, and r/Veterans the highest-error communities, while r/AskMen had the smallest total error.r/AskMen had the lowest over-estimation at 27% but the highest under-estimation at 12%.
- LLM alignment: 2–8 percentage-point F1 gains from contextualization were modest and concentrated in smaller instruction-tuned models, while larger Qwen3 showed negligible gains.Olmo-3-7B Instruct improved from 0.48 to 0.55 and Ministral-3-8B Instruct from 0.52 to 0.59; contextualization also risked shifting errors toward under-estimation.
- Interpretive caveat: Exploratory rationale themes from Olmo-3-7B-Think and Qwen3-30B-A3B should not be treated as faithful accounts of internal model decision processes.The rationale analysis was conducted to generate hypotheses about how conditioning redistributes errors.
5 Discussion and Conclusion
The study finds modest, uneven human in-group alignment and systematic model miscalibration, while emphasizing limitations in sampling, identity-context separation, and generalizability. Open-weight models often over-estimate low distress, and identity conditioning narrows this gap while risking under-estimation.
- Human judgment: OR = 1.18: contextualized in-group raters were modestly more likely to match their community’s distress aggregate, with significant heterogeneity across communities.The familiarity decomposition rules out depth of subreddit familiarity but cannot separate subreddit disclosure from in-group identity.
- Human judgment: Out-group raters labeled support-seeking far more often than veterans themselves, contrary to a simple insider-sensitivity account.The authors relate this divergence to military norms of self-reliance and stoicism.
- Model alignment: 23–42% versus 4–15%: open-weight models over-estimated distress on many more posts than they under-estimated, unlike the balanced out-group human error profile.On none-to-mild posts, random-stratum validation found 54% accuracy versus 44% on the full sample.
- Model alignment: Identity-based prompting narrowed the performance gap but shifted errors from over-estimation toward under-estimation.Under-estimation in r/TwoXChromosomes rose from 8% to 17%, so persona design alone is insufficient.
- Implications: Mental-health deployments should require calibration, human oversight, and community-specific validation rather than treating model severity judgments as calibrated point estimates.The authors identify fine-tuning on community data, community rationales, and judgment distributions as directions for improving alignment.
- Limitations: The study does not operationalize clinical distress, limiting direct comparability with clinical screening tools.Its purposeful sample also does not reflect natural base rates, and its in-group/out-group framing confounds identity with source-community disclosure.
- Limitations: Identity-group diversity, subreddit membership, and English-speaking US crowd-worker sampling constrain generalizability beyond the studied communities and raters.The authors caution that subreddit norms may not generalize to broader demographic groups and that within-group heterogeneity is not captured.
- Conclusion: The conclusion is that current LLMs are not drop-in substitutes for community judgment of psychological distress.Frontier systems differ in overall error direction, so deployment cannot assume uniformly inflated or conservative behavior.
Ethics Checklist
The checklist documents ethical review of the study’s assumptions, methods, data practices, participant protections, reproducibility, and potential societal impacts.
- Research design and limitations: The paper reports that its study design and analytic approach are appropriate for the research questions and that limitations and potential biases are discussed.The checklist references purposeful sampling, US-centric annotations, and within-group heterogeneity as potential artifacts.
- Reproducibility and data sharing: The paper reports documentation and release plans for prompts, codebooks, code, and data, while noting that platform policies restrict sharing curated post content.Only public post IDs can be shared without the original content, and a dataset datasheet is included.
- Impacts and misuse: The checklist states that the paper discusses misclassification costs, negative societal impacts, and potential misuse in mental-health deployments.The discussion distinguishes unnecessary interventions from missed genuine distress and addresses risks related to over-estimation and equitable deployment.
- Human subjects and privacy: The authors describe informed consent, participant compensation, content warnings, and procedures to reduce privacy risks in Reddit excerpts.Reddit posts were publicly available, annotators consented through Prolific, and excerpts were paraphrased with identifying details altered.
A1 Example Posts
Table A1 presents illustrative posts from four communities together with aggregated in-group, out-group, and LLM judgments for distress and support-seeking.
- Illustrative examples: The examples show that in-group, out-group, and LLM perspectives can diverge when evaluating the same post.The table covers both distress severity and support-seeking judgments across four communities.
A2 Codebook
The codebook was developed through iterative qualitative coding, repeated reliability checks, and collaborative discussion of disagreements and ambiguities.
- Codebook development: The research team collaboratively refined operational definitions of distress and support-seeking through multiple coding rounds.The first author and two undergraduate research assistants met twice weekly to adjust category definitions and refine the codebook.
A3 Annotation task
The annotation task recruited demographically screened Prolific participants, assigned them to identity-based in-group or out-group arms, and collected randomized post judgments after consent and comprehension checks.
- Participant screening: Participants were divided into 14 demographic rater groups, including six in-group groups aligned with the targeted Reddit identities.The six in-group identities were men, women, non-binary individuals, veterans, mothers, and fathers.
- Sampling arms: The study balanced community arms around 800 in-group or out-group ratings, with flexible out-group allocations for the two parenting communities.The parenting allocations varied between non-target-gender and non-parent raters while preserving 800 out-group ratings per parenting subreddit.
- Study flow: Eligible participants consented, passed comprehension checks, received contextualization when assigned as in-group raters, and were randomly assigned 30-post bundles.The flow also included demographic filtering, subreddit familiarization for applicable in-group raters, and shuffled post queues.
- Quality assurance: Quality assurance used codebook review, a ten-post practice round, and reference-label checks before the main annotation task.The supplied passage states that these safeguards were intended to support attentive human annotations.
- Example materials: The task included example posts from mothers, fathers, and non-binary participants illustrating the kinds of community discussions being judged.The examples describe parenting stress, an autistic child’s behavior, and coming out to parents.
N M M+
The study illustrates aggregated distress and support-seeking judgments while documenting safeguards and retention criteria intended to improve annotation reliability.
- Table A1 presents shortened Reddit examples alongside aggregated IG, OG, and LLM judgments for distress and support-seeking.
- Three attention-check items were embedded in each 30-post bundle, and participants failing two were excluded.
- AI-use safeguards prohibited participant assistance, blocked known browser agents, and flagged text selection and copying.
- The authors acknowledge that no protocol can fully eliminate participants’ use of AI during participation.
- After exclusions, posts required at least four IG and four OG ratings; two posts were dropped, leaving 9,587 judgments.
A4 Post-task survey
After annotation, participants reported their perceptions of task demands, effort, instruction clarity, and—among IG raters—understanding of subreddit norms.
- Participants completed a short post-task survey eliciting their perceptions of the study.
- Agreement was measured on a 5-point Likert scale from Strongly Disagree to Strongly Agree.
- Survey statements assessed whether the task was mentally demanding and required substantial effort.
- Participants rated whether the labeling instructions were clear.
- IG participants additionally rated their confidence in understanding the subreddit’s core norms.
5. Which part of the label was hardest?
Participants identified which aspect of the labeling task they found hardest, including distress severity, support-seeking, both, neither, or overall ease.
- Respondents chose among distress severity, support-seeking, both equally, neither, or finding the task easy overall.
6. Did any post make you feel distressed or uncomfortable?
In-group and out-group raters diverged because community context shaped how they interpreted masked communication, baseline experiences, contextual stressors, and the timing of distress. The analysis identifies recurring disagreement patterns in both human and LLM judgments, including salient-phrase anchoring and exemplar-driven normalization.
- Decoding Masked Communication: Community insiders recognized indirect, minimized, humorous, or culturally embedded expressions that outsiders often interpreted literally.Veteran in-group raters described reading hidden feelings between the lines, while humor could both signal shared coping and complicate interpretation.
- Community Baseline Calibration: Community baseline knowledge led insiders to normalize recurring shared experiences that outsiders sometimes pathologized as more severe distress.Examples included fireworks triggers among veterans, period dysphoria among non-binary people, and chronic sleep deprivation among parents.
- Contextual Integration: Insiders integrated systemic stressors, relationship dynamics, impairment, and protective context, whereas outsiders more often focused on surface language or isolated burdens.This produced contrasting ratings for cumulative financial pressures and family boundary violations that insiders interpreted as threats to stability or reactivated trauma.
- Temporal Anchoring: Insiders generally prioritized current functioning, while outsiders often anchored on historical trauma even when posts described recovery or present stability.The groups differed on veteran recovery narratives, although both acknowledged ambiguity when posts used past-tense descriptions.
- LLM Rationale Patterns: LLM rationales showed salient-phrase anchoring and impairment over-inference in over-estimation, while exemplar-driven normalization was associated with conditioning-related corrections and occasional over-correction.The taxonomy links model rationale patterns to structured interpretive capacities also distinguishing in-group from out-group human readers.