Source-linked AI summary

SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue

Stephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang, Yaling Shen, Hiu Weh Naomi Chow, Heung Ying Lai, Xiangyu Zhao, Qingyang Xu, Zhongxing Xu, Jiahe Liu, Guilherme C. Oliveira, Vincent Lee, Zongyuan Ge, Dominic Dwyer

arXiv:2609.01548v1cs.CL

TL;DR

Existing evaluations provide limited evidence about how LLMs recognise and respond to stigma in conversational and group settings. SDARE-Bench addresses this gap with a scenario-based benchmark covering stigma detection and open-ended responses across dyadic and multi-speaker dialogue. Across evaluations, models showed recurring stigma-related failures, with especially elevated expression under constructed group pressure.

  • Problem

    Existing stigma evaluations are scarce and often use static or fixed-format tasks, while group dynamics and conversational response behaviour remain underexamined.

  • Method

    SDARE-Bench evaluates stigma detection and open-ended response generation in dyadic and group conversations using psychologically grounded, expert-curated scenarios and an annotated response classifier.

  • Results

    Models often failed to identify stigma components and showed weaker, more stigma-reinforcing responses in socially complex settings, with group pressure raising stigma expression from 79.9% to 97.5%.

  • Takeaways & Limitations

    Stigma response is a recurring LLM safety vulnerability that requires evaluation and mitigation attentive to multi-speaker interaction and social dynamics.

  • Takeaways & Limitations

    The benchmark focuses on English text-only interactions and does not isolate the effects of speaker number, context, turn-taking, and distributed social cues.

Abstract

from arXiv · show

Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound effects on people and communities, benchmarks remain scarce. Existing general-domain evaluations typically rely on static prompts and fixed-format tasks, overlooking conversational contexts and audience effects in everyday communication. To address these gaps, we introduce SDARE-Bench, the first scenario-based benchmark evaluating both stigma detection and open-ended response generation in LLMs, comprising 1,138 dyadic queries and 1,388 group dialogue. Empirical results across 8 LLMs consistently demonstrate poor identification of stigma components, especially in group dialogues. In open-ended response generation, stigma expression was substantially higher in group settings than in dyadic, with weaker resistance to stigma and more unrealistic advice. Responses were evaluated using a classifier trained on 1,392 human annotated responses. In constructed group pressure settings, stigma expression rates further increased to a striking average of 97.5%. Our findings identify stigma response as a recurring LLM safety vulnerability, especially in socially complex conversational contexts.

1 Introduction

Stigma can produce status loss, exclusion, and unequal access, while LLM use creates opportunities to reproduce or amplify it. Existing evaluations provide limited insight into conversational stigma recognition and response, motivating SDARE-Bench’s dyadic and group settings.

  • Stigma involves negative social attributions that can lead to status loss, devaluation, exclusion, and inequitable access to essential resources.Psychological accounts distinguish stereotypes, prejudice, and discrimination as core stigma components.
  • LLMs are increasingly used in decision-making and support contexts where they may establish or amplify stigma.Prior evidence indicates negative associations toward stigmatised groups and differential recommendations with negative downstream consequences.
  • Existing evaluations focus largely on overt harmful content, demographic bias, static prompts, and fixed-response formats, leaving conversational stigma underexamined.Dedicated stigma benchmarks remain scarce and often use masked prompts, short vignettes, multiple-choice questions, or Likert ratings.
  • Three risks remain under-addressed: indirect stigma can evade harmful-content detection, fixed formats reveal little about appropriate responses, and group dynamics are largely omitted.Group interactions can reinforce, normalise, challenge, or resist stigmatising statements.
  • SDARE-Bench evaluates stigma detection, underlying components, and open-ended assistant responses across dyadic and multi-speaker settings.The benchmark is grounded in psychological literature and covers 93 stigma types, four stigma sources, and interactional speaker roles.
  • The benchmark combines scenario-based dyadic queries and group dialogues with expert-annotated response evaluation to expose failures beyond fixed-format testing.Its stated contributions include social-dynamics coverage, dual detection and generation evaluation, and scalable analysis of stigma responses.

2 Related Work

Prior stigma and bias benchmarks establish that LLMs can encode negative associations, but commonly restrict context, response format, or domain. SDARE-Bench broadens evaluation to socially grounded, open-ended dyadic and group conversations.

  • Existing stigma benchmarks use masked-token prompts, static sentences, handwritten templates, or constrained yes/no/can’t-tell responses.These designs quantify stigma or social decisions but provide restricted conversational and response contexts.
  • Mental-health stigma studies similarly rely on fixed vignettes, multiple-choice social-distance or danger ratings, and classification-oriented analyses.Recent work also examines how stigma features and guardrail filters shape model outputs.
  • Across prior benchmarks, restricted conversational context, closed response formats, or domain-specific settings limit the evaluation of socially grounded stigma responses.SDARE-Bench addresses this by testing open-ended generation in scenario-based dyadic and multi-speaker conversations across broader domains.
  • General harmful-content benchmarks target toxicity, offensive language, hate speech, maladaptive content, or unethical content, whereas stigma often operates indirectly.Indirect expression creates a distinction between stigma evaluation and overtly harmful-content detection.
  • Common bias benchmarks study demographic categories such as race, gender, and religion, while SDARE-Bench shifts the target toward broader stigma phenomena.The related-work comparison frames stigma as extending beyond conventional demographic bias categories.
  • SDARE-Bench covers 93 stigma types and includes discriminatory behaviours such as avoidance, coercive treatment, and withholding help.These behaviours are described as central to stigma theory but largely absent from existing bias benchmarks.

3 SDARE-Bench

SDARE-Bench evaluates whether models detect stigma and respond appropriately in English dyadic and group conversations. Its design combines psychological operationalisation, expert-guided scenario generation, and quality control.

  • SDARE-Bench evaluates stigma detection and appropriate conversational responses in both dyadic queries and group dialogues.The design enables evaluation of audience effects and group pressure.
  • The benchmark operationalises stigma through established psychological constructs rather than treating it as a single undifferentiated label.Its first design principle is grounded stigma operationalisation.
  • Expert-in-the-loop, schema-guided generation produces realistic socially situated items while maintaining coverage across labels.This is the benchmark’s second design principle.
  • Expert-centred quality control filters harmful or low-quality outputs before benchmark inclusion.Quality control is the third stated design principle.
  • SDARE-Bench organises stigma using five structured components, with definitions supplied in an appendix.The supplied passage introduces the component-based schema without enumerating all five components.

I. Social Scenario Selection

The benchmark constructs socially situated stigma scenarios by drawing on everyday activities, selecting plausible contexts and stigma types, and representing stigma sources and components through structured labels.

  • Social Scenario Selection: SDARE-Bench draws on 1,997 everyday activities from the American Time Use Survey Activity Lexicon to construct realistic interpersonal contexts.The activity lexicon provides the starting pool for scenario selection.
  • Social Scenario Selection: Three judges rated activities for plausibility, retaining 188 dyadic and 127 group scenario contexts for stigma curation.The judges were GPT-5-mini, Claude-Haiku-4.5, and Gemini-2.5-Flash.
  • Stigma Type Selection: For each scenario, three judges independently ranked plausible stigma types from a 93-category taxonomy, retaining the top five suitable types.This broadens stigma coverage beyond mental-health-focused open-ended evaluations.
  • Stigma Source: The benchmark distinguishes public, self, structural, and associational stigma according to where stigma originates in the interaction or social environment.These four sources form one structured dimension of the benchmark schema.
  • Stigma Components: Stigma expression is represented through stereotypes, prejudice, and discrimination, covering beliefs, affective reactions, and discriminatory behaviour.The supplied passage explicitly introduces these three components and begins detailing their meanings.

V. Conversational roles

SDARE-Bench models stigma as a socially distributed phenomenon by varying conversational roles and configurations across dyadic and group dialogues. Its schema-guided generation and quality-control process supports realistic, label-aligned items across these settings.

  • Conversational roles: Five stigma-present roles—stigmatiser, target, reinforcer, defender, and bystander—represent how stigma is distributed across speakers.Standard dyadic items use stigmatiser, target, or reinforcer roles, while standard four-speaker group dialogues use stigmatiser, target, reinforcer, and defender roles.
  • Conversational roles: Self-stigma variants merge the target and stigmatiser into a target-stigmatiser across dyadic and group settings.In group dialogues, the remaining speakers take reinforcer, defender, and bystander roles.
  • Quality control: Expert consultation iteratively refined schemas and prompts until pilot outputs were plausible, stigma-label aligned, and free of overt cues.An anthropologist and psychologist with expertise in social and digital harms informed the refinement process.
  • Schema-guided generation: Schemas independently sample stigma sources, components, and conversational roles to counter bias toward milder labels and against more severe labels.Scenario and stigma type are selected from ranked candidates, while other components are sampled from uniform distributions.
  • Query and dialogue generation: Naturalistic utterances embed stigma through indirect framing and implicit assumptions, while matched controls use the same schemas without stigmatising labels.Single-turn schemas become user queries, whereas group schemas become four-speaker, eight-turn dialogues ending with advice or worry.
  • Quality control: Sequential quality control combined harmful-content filtering, human calibration, LLM judging, and threshold-based pairwise selection.GPT and Gemini outputs were compared during quality control before final item selection.

I. Harmful Content Filtering

The benchmark filters generated content for harmfulness and calibrates scalable review against expert judgments. This process uses expert ratings to select an automated judge for evaluating generated items.

  • Harmful content filtering: Generated items were removed if flagged by any of five safety models, excluding overtly harmful content from the stigma-focused benchmark.The filter removed 46 dyadic and 21 group items.
  • Human expert evaluation: Two psychologists with over 25 years of clinical experience independently rated a 100-item calibration subset using a 0–2 rubric.The subset contained 50 dyadic and 50 group items and assessed general quality and stigma-label alignment.
  • Scaled LLM review: Mean absolute error compares expert scores with LLM judge scores across items.In the equation, x_i denotes the expert score and x̂_i denotes the LLM judge score for item i.
  • Scaled LLM review: Expert agreement was high at MAE = 0.15, supporting reliable application of the quality rubric.The calibration subset was used to compare candidate LLM judges against expert ratings.
  • Scaled LLM review: Llama-3.1-70B-Instruct matched expert ratings more closely than Claude-Sonnet-4.5, with MAE = 0.27 versus MAE = 0.36.Llama-3.1-70B-Instruct was therefore selected as the automated judge for all generated items.

IV. Pairwise Selection and Quality Control

Pairwise quality control retained stigma-present items that preserved stigma alignment and controls with zero stigma presence, while applying rubric thresholds to other dimensions. The resulting benchmark contains 1,138 queries and 1,388 group dialogues.

  • Pairwise selection: Stigma-present items required non-zero stigma alignment, whereas stigma-absent controls required zero stigma presence.This retained both subtle and explicit stigma in positive items while excluding stigma from controls.
  • Pairwise selection: All other rubric dimensions had to reach at least 85% of the maximum possible score, removing 3 dyadic and 117 group items.Filtering used Llama-3.1-70B-Instruct ratings after pairwise comparison of GPT and Gemini outputs.
  • Final benchmark composition: 1,138 queries and 1,388 group dialogues comprise the final SDARE-Bench.Full label distributions are reported in Appendix E.

4 Experiments

The experiments evaluate stigma detection and open-ended response generation across dyadic and group conversations, finding that conversational complexity worsens component classification and increases response failures. Group pressure further elevates stigma expression, while open-ended evaluation reveals failures that fixed-format benchmarks may miss.

  • Task I: Stigma Detection: Models detected stigma presence generally more accurately in group dialogues but classified stereotype, prejudice, and discrimination less accurately.The richer context increased stigma salience without enabling more precise characterisation of stigma components.
  • Task I: Stigma Detection: DeepSeek-V3.1 achieved the strongest overall detection performance, whereas Qwen3-8B and Mistral-7B showed the weakest overall performance.The weaker models were particularly limited on finer-grained component classifications.
  • Task I: Stigma Detection: Explicit label definitions changed most results modestly, while dyadic source and role accuracy increased by 7.9 and 6.0 pp and group role accuracy decreased by 10.0 pp.The overall conclusions remained unchanged after re-evaluation with definitions.
  • Task II: Stigma Response: Group dialogues had higher stigma, stereotype, prejudice, discrimination, unrealistic-advice, and quality-issue rates, alongside uniformly lower active pushback against stigma.Overly generalised advice showed no consistent pattern across models.
  • Task II: Stigma Response: 97.5% mean stigma expression under group pressure exceeded 79.9% in standard group stigma settings, with group pressure increasing the odds by a factor of 12.0.The difference remained after controlling for source model, format, input length, generation model, scenario, and stigma type fixed effects.

5 Discussion

SDARE-Bench reveals that stigma failures extend beyond binary detection: models struggled with components and open-ended responses, especially in group dialogue. Group settings also increased stigma reproduction and exposed stronger alignment with socially dominant framings.

  • Models often detected stigma presence but underperformed at identifying its source and components, showing why evaluation must extend beyond binary detection.
  • Models performed worse in group dialogue than dyadic queries across nearly all measured dimensions for detection and response.The authors associate this gap with longer context, multiple speakers, and distributed stigma cues.
  • Group-facing applications may require stigma-aware safeguards because socially complex settings produced more consistent model failures.The paper names collaborative tools, workplace chats, and clinical decision support as example settings.
  • Models were more likely to reinforce stigma when users acted as stigmatisers or reinforcers, and less likely when users were targets or self-stigmatising speakers.
  • Replacing the target and defender with additional reinforcers under group pressure significantly increased stigma reproduction.The authors connect this pattern to apparent prioritisation of user or group perspectives over protecting stigmatised speakers.

6 Conclusion

The paper introduces SDARE-Bench, a scenario-based benchmark for stigma detection and response across dyadic and multi-speaker conversations. It shows that models miss stigma components, reinforce or introduce stigma, and give unrealistic advice, particularly in multi-speaker and group-pressure settings.

  • SDARE-Bench is a scenario-based benchmark evaluating LLM stigma detection and response across dyadic and multi-speaker settings grounded in psychological literature.
  • Models struggled to recover stigma components, reinforced and introduced stigma during open-ended generation, and gave unrealistic advice.
  • These failures occurred more frequently when the stigmatised target was absent and in multi-speaker dialogues involving group pressure.

Limitations

SDARE-Bench is limited to English text-only interactions, does not isolate the effects of different group-dialogue factors, and evaluates reinforcement or resistance without prescribing mitigation methods or ideal responses.

  • The benchmark focuses on English text-only interactions, so multilingual and cross-cultural extensions are needed for broader coverage.
  • Its dyadic-versus-group comparison does not isolate speaker number, context, turn-taking, or distributed social cues.
  • The response evaluation measures whether models reinforce or resist stigma but does not develop mitigation methods or prescribe ideal responses.

Ethical Considerations

The benchmark uses synthetic data, expert consultation, licensing controls, and manual review, while restricting use to research and safety evaluation. It also cautions that its English-language results are culturally bounded and unsuitable for deployment decisions or universal stigma-safety claims.

  • SDARE-Bench contains no real user conversations or personal data, and external resources were used under their documented licenses and terms.
  • Three psychologists and one anthropologist voluntarily advised prompt design, evaluation metrics, and output assessment with informed consent and no identifiable information collected.
  • The benchmark is intended strictly for research and safety evaluation, not real-world decision making, profiling, or uncontrolled generation of stigmatising content.
  • Because stigma varies across languages, communities, and social norms, this English-language benchmark should not be treated as a universal measure of stigma safety.
  • ChatGPT assisted with code debugging, grammatical refinement, and icon generation, while authors manually reviewed and verified all outputs.

E Distribution of Stigma Related Labels in Benchmark Questions

The benchmark distributes stigma-related labels across individual and group dialogue settings, with examples spanning stigma types, sources, social roles, and stereotype, prejudice, and discrimination dimensions. The supplied passages show schema-level label assignments and dialogue roles, but not the numerical distributions reported in the referenced tables.

  • The benchmark records stigma-label distributions separately for individual and group dialogues.The supplied table references identify the distribution tables but provide no numerical frequencies.
  • Examples assign stigma types such as Speech Disability, Old Age, Drug Dependency Remitted, and Homeless across dyadic and group scenarios.
  • Group dialogue schemas additionally encode interactional roles such as stigmatiser, target, reinforcer, and defender.
  • The illustrated group scenario links housing instability with assumptions about repayment reliability and exclusion from a pooled loan.
  • Label schemas pair stigma sources with stereotype, prejudice, and discrimination labels, including public/incompetence/contempt/avoidance for a group borrowing scenario.
Loading 2609.01548v1…