Source-linked AI summary
How Mental Health Self-Disclosure Becomes Visible: Evidence from Eight Conditions on Reddit
Renkai Ma, Lingyao Li, Shanting Chen, Chen Chen, Fan Yang, Yuanyuan Lei
TL;DR
Research has not established how mental-health language becomes visible around self-disclosure across conditions or whether Reddit engagement tracks it. The study aligns 89,605 posts from 739 users across eight conditions to removed disclosure anchors and measures surrounding language and engagement. Burden visibility is uneven and persistent, while corrected language–engagement associations are sparse, supporting disclosure as a waypoint rather than a complete visibility event.
Problem
How mental-health language becomes visible around self-disclosure across conditions, and whether community engagement tracks it, remain underexamined.
Method
The study removes one diagnosis disclosure per user, aligns surrounding Reddit posts to that anchor, and analyzes language-visible burden, themes, and engagement associations.
Results
Across eight conditions, burden peaks immediately before disclosure for six conditions, varies in timing for PTSD and BPD, remains visible afterward, and only 9 of 360 correlations survive correction.
Takeaways & Limitations
Disclosure is an unevenly visible process, and recorded Reddit engagement is not a reliable indicator of clinical need within these cohorts.
Takeaways & Limitations
The selected anchor is not necessarily the author’s first disclosure or diagnosis date, and Reddit users are not representative of all people with mental-health conditions.
Abstract
from arXiv · showhide
People share mental health diagnoses on social media, yet how such language becomes visible around their self-disclosure, and whether community engagement tracks it, remain unexamined across conditions. We analyze 89,605 Reddit posts from 739 users across eight conditions, removing each user's diagnosis disclosure and aligning their surrounding posts to that anchor. Within the pre-disclosure year, language-visible burden was highest in the month before disclosure for six conditions, earlier for post-traumatic stress disorder and furthest from it for borderline personality disorder, and remained visible afterward rather than resolving. The theme Seeking Clinical Explanations showed the largest early-to-late difference before disclosure in five conditions, yet engagement rarely tracked what users wrote: only 9 of 360 language--engagement correlations survived correction. Disclosure is therefore a waypoint in an unevenly visible process, and we offer implications for community practice and platform design where engagement metrics do not reflect clinical need.
1 Introduction
The study examines how mental-health language becomes visible around self-disclosure and whether community engagement reflects that language across eight conditions. Using event-time analysis, it finds uneven, persistent burden visibility and sparse alignment between language and engagement.
- Prior work treated disclosure-adjacent language as predictive input or examined audience response, leaving the process connecting distress, diagnosis-related sense-making, and engagement less studied.
- The study analyzes 89,605 Reddit posts from 739 users across eight conditions, aligning surrounding posts to each user’s diagnosis disclosure.
- Language-visible burden peaks in the month before disclosure for six conditions, earlier for PTSD, and furthest from disclosure for BPD, then remains visible afterward.
- Seeking Clinical Explanations has the largest early-to-late pre-disclosure difference for five conditions.
- Only 9 of 360 language–engagement associations survive correction, indicating that recorded engagement rarely tracks condition-relevant language.
2 Related Work
Related work shows that self-disclosure supports cohort construction and research on language, support, and audience response, but typically treats disclosure cross-sectionally or within single conditions. This study addresses the missing common temporal design while distinguishing expressed language from observable engagement.
- Self-reported diagnosis statements have been used to construct condition-based cohorts, but prior analyses often did not separate posts before and after disclosure.
- Depression research provides a temporal precedent, finding pre-disclosure increases in anxiety, sadness, and cognitive processing that later declined toward matched-control levels.
- Prior studies across anxiety, PTSD, OCD, bipolar disorder, ADHD, autism, and BPD examined context, support, discourse, or cohorts without a shared disclosure-centered timeline.
- No prior study integrated condition-specific language, peer support, and audience response within a common temporal design.
- The paper combines the Disclosure Processes Model with networked visibility to distinguish what authors express from how content becomes socially observable.
3 Methods
The study constructs a rule-based, multi-condition Reddit cohort around one selected diagnosis disclosure per author, removes that anchor, and analyzes surrounding posts in fixed event-time windows. It measures condition-specific language and tests its co-occurrence with engagement using corrected statistical analyses.
- The corpus was built by identifying candidate disclosure authors, retrieving posting histories, and screening those histories to establish the cohort.
- The final cohort contains 739 authors and 89,605 surrounding posts after rule-based anchor selection and removal of one disclosure post per author.
- Posts were assigned to eight windows spanning the year before and after the anchor, from far-pre to far-post.
- Condition-specific burden dimensions were detected through predefined search terms applied only to the condition named in each disclosure anchor.
- Primary-theme hits were averaged within author and event-time windows, allowing posts to receive multiple non-mutually-exclusive theme hits.
- RQ3 tested 360 Pearson language–engagement correlations across 15 indicators, three engagement fields, and eight conditions, applying Benjamini–Hochberg correction.
4 Results
Across eight conditions, language-visible burden followed uneven timelines around disclosure, with condition-specific content and limited alignment between language indicators and recorded engagement. Burden often remained visible after disclosure, while seeking clinical explanations was the most consistent late-pre shift before the anchor.
- RQ1: Timing: Six conditions peaked in the immediate-pre window, PTSD peaked earlier in near-pre, and BPD peaked furthest from disclosure in far-pre.The immediate-pre maximum applied to ADHD, autism, depression, OCD, bipolar disorder, and anxiety.
- RQ1: Timing: 0.049 was the largest descriptive pre-disclosure burden mean, compared with 0.009 for ADHD; BPD averaged 4.9% dimension-share in far-pre.These maxima differed in both magnitude and temporal location across conditions.
- RQ1: Timing: None of eight paired tests had p<.05 for far-pre versus immediate-pre burden, so within-author comparisons did not support a common increase toward disclosure.The smallest unadjusted p-value was .053 for ADHD, with a mean difference of 0.007.
- RQ1: Timing: Burden remained visible after disclosure: ADHD and autism peaked in immediate-post, while bipolar disorder and anxiety peaked in near-post.The other four conditions retained their pre-disclosure maxima, and none of eight immediate-pre versus immediate-post paired tests had p<.05.
- RQ2: Content: The leading late-pre burden dimension differed across every condition, including Organizational Lapses for ADHD, Social Camouflaging for autism, and Reassurance Seeking for OCD.The composite burden score therefore represented condition-specific forms rather than one common burden profile.
- RQ2: Content: Seeking Clinical Explanations had the largest early-to-late pre-disclosure difference in five conditions, while other conditions showed distinct leading themes.Examples include a 0.106 difference for OCD’s checking and reassurance-seeking theme and −0.027 for depression’s personal-defectiveness interpretation theme.
- RQ3: Community Reaction: 9 of 360 same-window language–engagement correlations survived Benjamini–Hochberg correction, with five for ADHD and four for BPD.No corrected association involved the burden dimension-share score; the retained correlations were condition- and engagement-field-specific, with six positive and three negative.
5 Discussion
The discussion frames disclosure as a waypoint in an ongoing, condition-specific process whose temporal, interpretive, and engagement visibility do not reliably indicate clinical status or need.
- Engagement-Bounded Visibility: The framework separates what authors express, what engagement records, and what can be inferred about clinical experience.Language indicators and platform reactions describe observable content and activity but cannot establish underlying experiences, intentions, empathy, or understanding.
- Temporal Visibility: Disclosure marks an ongoing sensemaking process rather than the onset of a condition, with burden language remaining visible throughout the post-disclosure year.Condition-relevant experiences can precede explicit naming and remain observable afterward.
- Temporal Visibility: Burden trajectories differed across conditions: peaks occurred immediately before disclosure for ADHD but during the far-pre period for BPD.Across the full observation period, four conditions reached their highest condition-window mean only after disclosure.
- Interpretive Visibility: Interpretive visibility was condition-specific, ranging from abandonment fear among BPD disclosers to reassurance-seeking among OCD disclosers.Seeking Clinical Explanations showed the largest pooled early-to-late pre-disclosure difference across five conditions.
- Engagement-Bounded Visibility: Only 9 of 360 same-window language–engagement associations survived correction, and none involved language-visible burden.The surviving associations varied across conditions and indicators, underscoring that engagement does not uniformly track expressed burden.
- Implications: Highly visible or highly engaged disclosures may overlook users who express distress indirectly or receive little community response.The framework therefore cautions against equating absent digital evidence with absent distress.
Practice and Social Media Platform Design
The paper translates heterogeneous disclosure trajectories and sparse engagement correspondence into community and platform practices that support varied expressions without treating visibility as clinical need.
- People who disclose: Across all eight conditions, condition-relevant burden language was measurable months before disclosure, peaking immediately before disclosure for six conditions, earlier for PTSD, and furthest from disclosure for BPD.Community resources can treat preceding posts as part of an ongoing sensemaking process.
- People who disclose: Low engagement provides limited information about whether authors’ experiences matter or have been understood.Users may underestimate how many people encounter their content, while platform mechanisms obscure who sees or responds.
- Peer supporters and community members: Effective peer support should respond to condition-specific experiences without assuming that those expressions establish a diagnosis.Leading burden dimensions differed across cohorts, including masking and sensory overload for autism, reassurance-seeking for OCD, abandonment fear for BPD, and forgetfulness and disorganization for ADHD.
- Peer supporters and community members: Structured check-in threads, norms encouraging replies to zerocomment posts, and dedicated peer-support formats could broaden response opportunities without requiring popularity first.The paper presents this as a plausible response to the possibility that gradual or less affectively intense difficulties receive fewer visible responses.
- Platform moderators, designers, and public-health practice: Engagement and burden should not be treated as unrelated or as reliable triage signals because cohort sizes and unmeasured exposure differences constrain inference.Platforms have no basis for assuming highly engaged posts reflect greater burden or less-engaged posts reflect less need.
- Platform moderators, designers, and public-health practice: Platforms could create chronological unanswered queues, structured check-in spaces, or optional moderator acknowledgments for posts with little initial engagement.Language indicators may support aggregate community monitoring only with ongoing validation, not individual pre-disclosure identification.
6 Limitations, Ethics, and Future Work
The findings are bounded by limitations in disclosure-anchor selection, available posting history, measurement validity, aggregation, representativeness, and engagement analysis. Ethical scope is also limited to analysis of public Reddit posts without interaction with authors.
- Scope and measurement: Disclosure anchors may miss earlier disclosures, are not diagnosis dates, and leave other disclosure posts in users’ timelines.The selected event is the earliest eligible same-condition post at or before the initial candidate, prioritizing precision over cohort size.
- Scope and measurement: 55 retained authors had fewer than five pre-disclosure posts, including 19 with none, limiting which event-time windows they contribute to.
- Scope and measurement: Lexical indicators may miss context, sarcasm, negation, and relevant language, and do not measure clinical severity or diagnostic thresholds.Human coder agreement was substantial rather than near-perfect, and fewer than half of the automatically identified items were validated.
- Inference boundaries: Condition-window means and pooled early- and late-period estimates do not establish within-author change, symptom onset, or effects caused by disclosure.
- Inference boundaries: Reddit users are not representative of all people with mental health conditions, while engagement correlations are unadjusted and do not establish prediction or causality.Same-window tests treat author-windows as independent, and lagged analyses cannot rule out confounding.
- Ethics: The analysis used public Reddit posts about sensitive mental health topics without interacting with the people who wrote them.
7 Conclusion
Across eight conditions, the study finds that language-visible burden and its timing around first-person diagnosis disclosure are condition-specific. Burden often peaks before disclosure, remains visible afterward, and shows limited alignment with recorded engagement.
- Conclusion: Across eight conditions, the timing and content of language-visible burden around first-person diagnosis disclosure are condition-specific.
- Conclusion: Most conditions show peak burden language in the month before disclosure, while burden remains visible across the post-disclosure year.
- Conclusion: Seeking Clinical Explanations shows the most consistent early-to-late pre-disclosure difference.
- Conclusion: Corrected associations between language indicators and Reddit engagement fields are sparse and limited to two conditions.
- Conclusion: The candidate-disclosure query combines 41 first-person disclosure phrases with 33 mental health terms as a recall-oriented first pass.The query cannot separate an author’s own diagnosis from quoted or hypothetical diagnoses, and BPD lacks condition-specific terms.
B Author-Aggregated Sensitivity Results
Author-level sensitivity analysis collapses observed event-time windows into one row per author before re-estimating correlations. Coefficient signs remain consistent, although most coefficients shrink after repeated within-author windows are removed.
- Author aggregation: Author-level sensitivity analysis averages each author’s observed event-time windows into one row before re-estimating Pearson correlations and BH-adjusted q-values.
- Author aggregation: Coefficient signs stay consistent across author-window and author-level analyses, while most coefficients shrink after repeated windows within authors are removed.
- Search-term reference: The search-term table marks the eight study conditions, with BPD captured through general and personality-disorder terms.
- Reported outputs: Table 4 reports the 9 of 360 same-window Pearson correlations meeting the BH-adjusted q≤.05 threshold across 15 language indicators.
- Reported outputs: Table 5 compares the nine corrected same-window correlations with author-level re-estimates after collapsing each author’s observed windows to one mean row.
- Measurement boundary: The analysis treats burden measures as language indicators rather than clinical measures and Reddit engagement fields as neither exposure nor support measures.
C.1 Event-Time Indexing
Event time indexes each surrounding post relative to its disclosure anchor in days, then assigns the post to fixed intervals. The appendix links these measurements to the manuscript’s reported outputs.
- Event-time indexing: Event time is the number of days between a surrounding post’s timestamp and the disclosure-anchor timestamp.
- Event-time indexing: The denominator converts seconds to days before posts are assigned to fixed event-time intervals.
- Event-time indexing: Each event-time interval includes its lower boundary but excludes its upper boundary.
- Measurement reference: The appendix measurement-to-output map links manuscript tables and visualizations to the equations and citations defining or calculating their measurements.
C.2 Operational Criteria for Disclosure-Anchor Selection
Disclosure anchors are selected through conservative binary eligibility rules, then chronologically fixed to the earliest eligible same-condition post within the available study history. An audit score orders ties but does not determine eligibility.
- Audit and tie-breaking: The condition-match score records inspectable evidence and orders candidates sharing the earliest timestamp, but no minimum score threshold determines eligibility.A post identifier breaks any remaining tie.
- Binary eligibility: Eligibility requires a qualifying first-person statement and identification of the assigned condition as the only named study condition.The binary gate excludes posts naming multiple study conditions.
- Binary eligibility: Uncertainty, denial, questioning, third-party attribution, subjectless diagnosis wording, and symptom-only wording disqualify a candidate disclosure.These rules set the first-person disclosure indicator to zero or exclude the candidate.
- Chronological selection: The selected anchor is the earliest eligible post for the author’s assigned condition within the available study history and temporal boundary.Posts naming other study conditions neither reassign the author nor serve as the selected anchor.
- Scope: The anchor is not a verified diagnosis date or the author’s first-ever disclosure.It is the earliest eligible same-condition post found within the available study history and stated temporal boundary.
C.4 Burden and Affect Indicators
The study measures condition-relevant burden, affect, and themes through lexical indicators applied to surrounding posts. These indicators describe language patterns rather than clinical severity, support quality, or manually interpreted themes.
- Burden indicators: A burden-dimension hit is binary, and the dimension-share score is the proportion of condition-specific dimensions with at least one lexical hit.Repeated terms within one dimension do not increase the score.
- Burden indicators: Burden dimensions are literature-informed, condition-specific lexical categories applied only to the condition named in each disclosure anchor.Depression uses five dimensions, while each other condition uses four.
- Burden indicators: The secondary burden phrase rate counts matched burden phrases per 100 word-like tokens, with the denominator constrained to at least one.Tokens may contain letters, numbers, underscores, and apostrophes.
- Interpretive boundaries: These lexical metrics cannot validate a clinical scale, and relief-or-support language identifies the author’s post rather than support received from other users.The burden dimensions and affect categories describe text, not clinical status or the quality of support.
- Affect indicators: Affect indicators count condition-specific lexical matches per 100 word-like tokens across seven shared categories.Reported analyses use negative-affect and positive-affect rates, sentiment balance, and affective arousal; sadness, anger, and relief-or-support remain auxiliary outputs.
- Theme indicators: Theme assignment marks a post when any associated case-insensitive lexical pattern appears, and a post may receive multiple theme hits.The implementation applies the codebook rather than reproducing manual thematic interpretation.
C.6 Author-Window and Interval Estimates
The analysis aggregates posts into author-window means to compare event-time patterns while reducing the influence of unusually active authors. Reported intervals and associations remain descriptive and observational, with multiple-comparison correction applied to declared test families.
- Author-window aggregation: Author-window aggregation averages each author’s post-level measure within condition and event-time window before forming condition-level series.This reduces the influence of unusually active authors in event-time comparisons.
- Interval estimates: Approximate 95% intervals use standard errors across observed author-window means and do not represent population intervals.They summarize variation in the retained author-window sample.
- Window comparisons: Adjacent-window comparisons pool author-window observations, so an author observed in both windows contributes two values.The denominator counts author-window observations rather than unique authors.
- Paired differences: Paired far-pre to immediate-pre and immediate-pre to immediate-post differences compare the same authors across windows as descriptive longitudinal contrasts, not causal disclosure effects.The paired formulation is author-level.
- Engagement associations: Engagement correlations use Pearson association between language indicators and Reddit engagement fields within condition, after coercing missing or invalid engagement values to zero.A secondary Mann–Whitney comparison separates author-windows by the presence or absence of nonnegative indicators.
- Sensitivity and lagged checks: Author-level sensitivity averages language and engagement across observed windows before re-estimating correlations, but remains observational.Lagged checks add temporal ordering while remaining vulnerable to unmeasured confounding and selective posting.
- Multiple-comparison correction: Headline same-window Pearson claims use Benjamini–Hochberg-adjusted q-values across a 360-test family, while lagged analyses use separate 216-test families.Mann–Whitney and author-level sensitivity families are adjusted separately.