Source-linked AI summary
Social Bias Frames: Reasoning about Social and Power Implications of Language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, Yejin Choi
TL;DR
Existing approaches often reduce implied social bias to toxicity labels, despite the importance of explaining pragmatic implications and power dynamics. The paper introduces Social Bias Frames and SBIC to structure these meanings and evaluates baseline models for recovering them. Models achieve 80% F1 on high-level bias categorization but struggle with detailed social-bias inferences, while SBIC’s language distribution motivates caution about dialect- and identity-based bias.
Problem
Prior toxicity-classification approaches do not adequately represent implied social biases, and detailed explanations are needed to understand why language may be harmful.
Method
The paper defines Social Bias Frames with categorical and free-text inferences, collects SBIC, and establishes baseline models for recovering frames from unstructured text.
Results
Models achieve 80% F1 for high-level unwanted-bias categorization but struggle to accurately decode detailed Social Bias Frames.
Takeaways & Limitations
The study supports combining structured pragmatic inference with commonsense reasoning to model social implications of language.
Takeaways & Limitations
SBIC is predominantly White-aligned English, with 78% of posts in White-aligned English and fewer than 10% showing indicators of African-American English.
Abstract
from arXiv · showhide
Warning: this paper contains content that may be offensive or upsetting. Language has the power to reinforce stereotypes and project social biases onto others. At the core of the challenge is that it is rarely what is stated explicitly, but rather the implied meanings, that frame people's judgments about others. For example, given a statement that "we shouldn't lower our standards to hire more women," most listeners will infer the implicature intended by the speaker -- that "women (candidates) are less qualified." Most semantic formalisms, to date, do not capture such pragmatic implications in which people express social biases and power differentials in language. We introduce Social Bias Frames, a new conceptual formalism that aims to model the pragmatic frames in which people project social biases and stereotypes onto others. In addition, we introduce the Social Bias Inference Corpus to support large-scale modelling and evaluation with 150k structured annotations of social media posts, covering over 34k implications about a thousand demographic groups. We then establish baseline approaches that learn to recover Social Bias Frames from unstructured text. We find that while state-of-the-art neural models are effective at high-level categorization of whether a given statement projects unwanted social bias (80% F1), they are not effective at spelling out more detailed explanations in terms of Social Bias Frames. Our study motivates future work that combines structured pragmatic inference with commonsense reasoning on social implications.
1 Introduction
The paper argues that social bias is often conveyed through implied meanings rather than explicit wording, motivating structured representations that explain those implications. It introduces Social Bias Frames and SBIC, then finds models classify high-level bias more effectively than they recover detailed explanations.
- Social biases often arise through implied meanings that shape judgments, so understanding them requires reasoning about intent, offensiveness, and power differentials.
- Simple toxicity classifications can discriminate against minority groups and provide less informative explanations of why statements are harmful.
- Social Bias Frames model pragmatic bias frames by combining hierarchical categories with free-text implicatures about referenced groups and implied stereotypes.
- SBIC provides over 150k structured annotations spanning over 34k implications about approximately a thousand demographic groups.
- 80% F1 was achieved for high-level categorization of unwanted social bias, while models struggled to decode detailed Social Bias Frames.
- The study motivates combining structured pragmatic inference with commonsense reasoning about social implications.
2 SOCIAL BIAS FRAMES Definition
Social Bias Frames represent social-bias implications through categorical judgments and free-text explanations. Their variables distinguish offensiveness, intent, sexual references, group targeting, targeted groups, implied stereotypes, and in-group language.
- The frame design combines categorical and free-text inferences, refined using social-science literature and annotator agreement from pilot studies.
- Offensiveness denotes a post’s overall rudeness, disrespect, or toxicity and is labeled yes, maybe, or no.
- Intent to offend captures the author’s perceived motivation to offend and is distinct from offensiveness, with four categorical answers.
- Lewd or sexual references form a potentially offensive subcategory with yes, maybe, or no labels.
- Group implications distinguish individual-only attacks from statements that target groups and invoke intergroup power dynamics.
- The framework records targeted groups and implied stereotypes as free-text answers, while in-group language captures whether the speaker may belong to the targeted group.
3 Collecting Nuanced Annotations
SBIC was built from diverse potentially biased online sources and a hierarchical crowdsourcing task that elicits structured and free-text social-bias inferences. The resulting corpus contains 150k inference tuples, but its coverage and annotations have important distributional and agreement characteristics.
- 3.1 Data Selection: SBIC draws posts from offensive Reddit communities, microaggression data, toxic-language Twitter datasets, and known hate communities.The sources include intentionally offensive subreddits, tweets selected from prior toxic or abusive-language datasets, and documented white-supremacist, neo-Nazi, or violence-inciting communities.
- 3.2 Annotation Task Design: The hierarchical task asks workers to assess offensiveness, intent, lewdness, group implications, targeted groups, stereotypes, and possible in-group speech.Workers write two to four stereotypes for each selected group and indicate whether the speaker may belong to a referenced minority group.
- 3.2 Annotation Task Design: Three annotations are collected per post from workers in the U.S. and Canada, with optional coarse-grained demographic information.The study also reports institutional review board approval.
- 3.2 Annotation Task Design: 82.4% pairwise agreement and Krippendorf’s α=0.45 were observed overall, while in-group speech had α=0.17 despite 94% agreement.Agreement was 80.2% for the exact targeted group, and lewdness had the strongest agreement at 94% with α=0.62.
- 3.3 SBIC Description: 150k structured inference tuples cover 34k free-text group-implication pairs in SBIC.The corpus statistics are summarized in Table 3.
- 3.3 SBIC Description: Gender-based, race-based, and culture-based biases are the most represented targeted-group categories in SBIC.The dataset is predominantly written in White-aligned English, comprising 78% of posts, with fewer than 10% showing indicators of African-American English.
4 Social Bias Inference
The study frames Social Bias Frame inference as a hybrid classification and language-generation task using GPT-based models. It evaluates categorical predictions and generated group and implied-statement inferences, including constrained decoding.
- Training: Models predict five categorical variables and generate targeted groups and implied statements from linearized Social Bias Frames.Training combines classification and language generation, with frame variables ordered hierarchically and represented as task-specific vocabulary items.
- Model: GPT-based models encode the input with causal transformer blocks and predict the next token through a vocabulary-sized softmax distribution.The models attend only to preceding tokens because GPT is a forward-only language model.
- Inference: At inference, models generate frame tokens sequentially using either greedy decoding or sampling, with sampling repeated to produce ten candidates.Generation stops when the [END] token is produced.
- Inference: Constrained decoding searches globally probable categorical assignments to reduce inconsistencies between generated text and categorical variables.The method recomputes categorical probabilities after greedy decoding and searches assignments based on the generated candidate.
- Evaluation: Classification is evaluated with positive-class precision, recall, and F1, while generation uses BLEU-2, RougeL, and WMD.BLEU-2 and RougeL reward word overlap, whereas lower WMD is better.
- Results: Constrained decoding only slightly improves predictions across the reported evaluation tables.The authors report this effect in Tables 4, 5, and 7.
5 Results
Models perform well on some high-level social-bias classifications, but detailed Social Bias Frame inference and generation remain challenging.
- Classification: Models perform well on higher-level variables such as offensiveness and lewdness classification.Lewdness may benefit from lexical matching, despite heavy class skew.
- Classification: Whether language targets a group is harder to predict, while in-group language is more challenging because of subtlety and few positive instances.Most models default to never predicting in-group language.
- Classification: 1.9% improvement in F1 score is obtained for in-group language by SBF-GPT2-gdy with constrained decoding.It is the only model predicting positive in-group values.
- Generation: No model outperforms the others across all metrics on targeted-group and implied-statement generation.Targeted-group generation is easier than implied-statement generation because its output space is smaller.
- Generation: Constrained decoding slightly increases RougeL but slightly decreases BLEU for SBF-GPT2-gdy.The trade-off appears in the generation metrics.
- Error analysis: Manual analysis finds that the model generates generic stereotypes instead of post-relevant implications, except when lexical overlap is high.The evaluation used manually selected development-set examples from SBF-GPT2-gdy-constr.
6 Related Work
Related work includes toxicity and bias detection, inference about social dynamics, and commonsense reasoning about participants and situations.
- Bias and toxicity detection: Most toxicity-detection datasets formulate the task as binary classification, while some use multiple binary toxicity-related variables.These approaches motivate moving beyond a single toxicity label.
- Bias and toxicity detection: Recent work addresses subtler language biases such as microaggressions and condescension, which overlap with Social Bias Frames but are more narrowly scoped.The comparison concerns the kinds of biases modeled.
- Inference about social dynamics: Prior studies infer power dynamics about specific entities in conversations or narratives, while commonsense work models participants’ mental states.These lines of work address related social inference problems.
7 Ethical Considerations
The paper emphasizes ethical risks in detecting harmful language, annotating abusive content, and reusing public-forum data.
- Risks in deployment: Deploying harmful-language detection requires attention to metrics, demographic fairness, censorship, and dialect-based racial bias.The paper also suggests pairing offensiveness detection with community standards or counterspeech.
- Risks in annotation: Annotating abusive content can cause acute stress, so the study limited daily workloads, paid $7–12, and provided crisis resources.These measures were intended to mitigate annotation-related harms.
- Data ethics: Researchers and practitioners are urged to respect the privacy of authors whose posts appear in SBIC.The concern arises from using data available on public forums for research.
8 Conclusion
The paper introduces Social Bias Frames and SBIC, evaluates pretrained-language-model baselines, and finds that detailed social-bias inference remains difficult.
- Conclusion: Social Bias Frames provide a structured commonsense formalism for representing biased implications of language.The frames combine categorical variables with free-text social-bias inferences.
- Conclusion: SBIC contains 150k annotations on social-media posts, and the paper establishes baselines using large pretrained language models.The dataset and baselines support modeling social-bias implications.
- Conclusion: Models classify offensiveness more easily than they generate relevant social-bias inferences, especially when implications have low lexical overlap with posts.The conclusion calls for more sophisticated models for Social Bias Frames inference.