Source-linked AI summary
Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus
Mohammad Omar Khursheed, Mandira Sawkar, Ashiqur R. KhudaBukhsh
TL;DR
Computational evidence on the arguments used in historical suffrage debates is limited, especially regarding the type of sexism deployed. The paper analyzes Hansard speeches with LLM-based stance and sexism classification grounded in the Ambivalent Sexism Inventory, finding distinct rhetorical modes across opposing positions. These patterns support the relevance of distinguishing benevolent from hostile sexism in legislative discourse, while the analysis remains bounded by keyword retrieval, model judgment, and House of Commons coverage.
Problem
Computational analysis of the actual arguments used in the well-documented suffrage movement is limited, particularly regarding whether opposing sides used different types of gendered reasoning.
Method
The study classifies 6,531 suffrage-related Hansard speeches with LLMs for stance toward women’s political representation and hostile or benevolent sexism using the Ambivalent Sexism Inventory.
Results
Opposing and supporting speakers used fundamentally different rhetorical strategies: opponents were direct in their contempt, while sexist supporters used praise that confined women and mapped onto benevolent sexism.
Takeaways & Limitations
The findings provide naturalistic evidence that hostile and benevolent sexism serve different rhetorical functions across political stances, and that sexism type is informative beyond its presence.
Takeaways & Limitations
The analysis may miss relevant speeches lacking its search terms, relies primarily on one LLM judge, and is restricted to the House of Commons.
Abstract
from arXiv · showhide
The language a legislature uses to debate women's rights, even in favour of them, encodes systematic patterns of sexism that persist across two centuries. In this work, we analyse 6,531 speeches over 200 years of UK parliamentary debate (Hansard, 1803-2005) by using large language models to classify a speaker's perspective towards women's suffrage and political representation, as well as analyse sexist speech in parliament from the lens of the Ambivalent Sexism Inventory. We also release this parliamentary dataset, an organized and metadata-enriched version of the publicly available Hansard Corpus optimized for computational social science research, with 6.7 million speeches across 1.2 million debates, with 89% gender-matching for speeches by MPs from the House of Commons. We find that 54% of speeches opposing women's representation contain sexist content, compared to 21% of speeches that are for the cause, and that the two sides use fundamentally different types of sexism: anti-suffrage rhetoric combines hostile and benevolent framing, while pro-suffrage sexism is overwhelmingly benevolent. Female MPs support women's political rights at 93% compared to 70% for male MPs, a gap that closes only after enfranchisement. Our findings are evidence that benevolent and hostile sexism are used in different rhetorical contexts in a manner consistent with the theory of Ambivalent Sexism.
1 Introduction
This paper addresses limited computational analysis of the arguments used in Britain’s suffrage debates by combining a large Hansard dataset with LLM-based stance and sexism classification. It examines whether opposing positions used different rhetorical strategies and releases an enriched corpus for computational social science.
- Motivation: Computational analysis of suffrage arguments remains limited despite historians’ extensive documentation of the movement.The paper asks whether proponents and opponents differed in their premises or in the types of gendered reasoning they deployed.
- Motivation: The study uses the Ambivalent Sexism Inventory to distinguish hostile sexism from benevolent sexism in historical political discourse.The framework treats hostile sexism as overt antagonism and benevolent sexism as flattering or chivalrous framing that casts women as frail and helpless.
- Research design: 6,531 suffrage-related speeches from 1803–2005 were drawn from the Hansard corpus and classified for stance and sexist content using LLMs.The classifications were validated against human annotations on a 300-speech validation set.
- Dataset: 6.7 million British parliamentary speeches were organized into a large-scale, gender-matched dataset for computational social science research.The source corpus contains 1,197,828 debates and 6,783,015 speeches across House of Lords and Commons proceedings.
- Dataset: 89.3% coverage was achieved when matching House of Commons speakers to gendered MP records, compared with 1.2% in the House of Lords.The analysis therefore restricts its substantive work to the House of Commons.
- Research design: A two-tier keyword search identified 6,531 speeches spanning 1809–2004 from the broader Hansard corpus.The filter used explicit suffrage terms and proximity between women-related and voting-related terms.
4 Taxonomy
The paper annotates suffrage-related speeches along two independent dimensions: stance toward women’s political participation and the presence and type of sexism. Its taxonomy separates hostile and benevolent forms, including paired subcategories across paternalism, gender differentiation, and heterosexuality.
- Taxonomy: Two annotation axes capture each speech’s stance toward women’s suffrage and its presence and type of sexism.This dual-axis design distinguishes the direction of a position from the gendered reasoning used to express it.
- Stance: Four mutually exclusive stance categories are For, Against, Both, and Irrelevant.Irrelevant functions as a filter for speeches that do not substantively address women’s political representation or public life.
- Sexism: Hostile and benevolent sexism are independent flags, so a speech may contain either form, both forms, or neither.Table 2 pairs the forms across paternalism, gender differentiation, and heterosexuality.
- Sexism: Hostile sexism expresses overtly negative attitudes toward women, especially those challenging male authority or traditional gender roles.Its examples include dominative paternalism, competitive gender differentiation, and heterosexual hostility.
- Sexism: Benevolent sexism superficially praises women while constraining them as pure, fragile, complementary, or dependent on male protection.Its corresponding domains are protective paternalism, complementary gender differentiation, and heterosexual intimacy.
5 Annotation
The paper combines human annotation and LLM classification to identify stance and sexism in historical parliamentary speeches, then validates the approach and reports longitudinal patterns in representation debates.
- Human annotation: 300 speeches were sampled across eras, with two annotators independently labeling relevance, stance, and hostile or benevolent sexism using surrounding debate context.Annotators resolved disagreements through discussion after independent coding.
- LLM classification: Claude Sonnet 4.6 classified stance in 6,531 keyword-retrieved speeches using up to five preceding and following speeches as context.The relevance definition included broader political representation arguments arising in franchise debates.
- Validation: κ = 0.711 measured Claude’s agreement with consensus human labels, compared with human–human agreement of κ = 0.644.For For, Against, and Irrelevant classes, F1 was at least 0.75; the rare Both class was under-detected.
- Stance results: 55% of 6,531 extracted speeches were irrelevant; among 2,942 relevant speeches, 74% were For, 19% Against, and 7% Both.Opposition steadily declined over time, with negligible opposition remaining after 1950.
- Stance results: Female MPs supported women’s representation in 93% of speeches, compared with 70% for male MPs, and gender remained a significant predictor after controlling for decade.Male support rose from 54% in 1870–1899 to 91% post-1950, converging with female-comparable levels only after suffrage was settled.
- Sexism results: 54% of Against speeches contained sexism, compared with 21% of For speeches, with anti-suffrage rhetoric combining hostile and benevolent framings while pro-suffrage sexism was mostly benevolent.Among sexist For speeches, 81% were benevolent-only; among sexist Against speeches, 37% were hostile-only, 19% benevolent-only, and 44% both.
- Sexism results: Hostile sexism fell from 60% of sexist speeches in 1870–1899 to 27–30% after 1929, while benevolent sexism remained at 74–83%.The authors interpret this pattern as declining acceptability of overt hostility alongside persistent benevolent framing.
8 Discussion
Parliamentary sexism varied by political stance: opponents used hostile and benevolent reasoning, while supporters’ sexism was predominantly benevolent. This pattern aligns with ambivalent-sexism theory and remains relevant to contemporary debates over participation and inequality.
- Discussion: Opponents used hostile and benevolent sexism, whereas sexist supporters primarily used paternalistic praise that constrained women’s roles.The contrast reflects distinct rhetorical strategies rather than merely opposing policy conclusions.
- Discussion: The study provides naturalistic evidence that hostile and benevolent sexism serve different rhetorical functions in parliamentary discourse.The analysis extends an ambivalent-sexism framework previously developed primarily in laboratory and survey settings.
- Discussion: Progressive policy positions can coexist with sexist reasoning that portrays women as fragile, morally pure, or requiring special accommodation.Such assumptions may be harder to identify and challenge when attached to supportive positions.
- Discussion: The documented rhetorical structure may apply to contemporary and international debates over who deserves full participation in public life.The authors identify legislative debate corpora as settings where discriminatory reasoning can be studied beyond its direction alone.
9 Conclusion
The conclusion shows that sexist reasoning in parliamentary debate differed systematically by stance and persisted across eras. The released dataset and computational approach support further large-scale study of historical political discourse.
- Conclusion: 1873 hostile sexism asserted male authority by denigrating women, while later pro-rights sexism idealized women through paternalistic praise that restricted their roles.These modes map onto the hostile/benevolent structure predicted by the Ambivalent Sexism Inventory.
- Conclusion: Distinct hostile and benevolent rhetorical modes remained separated by speaker stance across contexts and eras.The authors interpret this separation as evidence that mechanisms of sexist reasoning are remarkably stable.
- Conclusion: The study examined over 6.7 million enriched Hansard speeches across 200 years, varying discourse analysis by stance, gender, and era.The dataset is released with reorganized speaker metadata for computational social science research.
11 Limitations
The study identifies methodological, representational, and interpretive limits affecting how its findings should be understood. These include model dependence, keyword-retrieval coverage, House of Commons scope, binary gender categories, and the descriptive nature of the framework.
- Model dependence: A single primary LLM judge was used, although cross-model validation suggests classifications reflect textual signal rather than model-specific artefacts.The additional models were GPT-5, Gemini 2.5 Flash, and DeepSeek V3.
- Retrieval coverage: Keyword retrieval can miss relevant speeches whose vocabulary diverges from the search terms, setting an upper bound on what the analysis can find.The relevance classifier filters false positives, but cannot recover relevant speeches missed at retrieval.
- Scope and detection: The analysis covers only the House of Commons, and detected sexism differences may partly reflect hostile sexism being easier for the LLM to identify than benevolent sexism.Both forms are predominantly negative-sentiment, but residual detection asymmetry remains possible.
- Gender representation: Binary gender classification reflects historical records rather than the full range of gender identities.A robustness check flipping 3% of speaker-gender labels preserved statistical significance in all 1,000 iterations.
- Interpretive scope: The framework is primarily descriptive and does not model historically contingent motivations for individuals’ beliefs.The analysis categorizes recurring expressions of gender bias rather than providing causal explanations.
A Suffrage Speech Extraction Keywords
Suffrage-related speeches were extracted with a two-tier, high-recall keyword search and then filtered for relevance. This design captures broad candidate material but introduces false positives and can miss relevant arguments using divergent vocabulary.
- Search design: The extraction pipeline used a two-tier keyword search over full speech text, with case-insensitive matching.Tier 1 targeted explicit suffrage terms; Tier 2 matched women or female near voting-related terms.
- Search design: 2,725 speeches matched explicit suffrage terms, while 3,806 matched proximity-based voting terms.The two tiers yielded 6,531 keyword-extracted speeches in total.
- Relevance filtering: 55% of extracted speeches were marked irrelevant by the downstream LLM classifier, filtering false positives such as male-franchise debates.The search was designed for high recall rather than precision.
- Retrieval limitations: The keyword search can miss relevant speeches when their vocabulary diverges from the retrieval terms.A cited example involved a women’s voting debate not captured because “vote” was not within 25 words of “women.”
- Retrieval limitations: Keyword matches also captured irrelevant passing mentions and debates about Irish representation, which annotators marked as irrelevant.These examples illustrate why downstream relevance classification was necessary.
- Future work: Future work could expand the keyword list, use more robust techniques, and scale the analysis with additional computational resources.
B Hansard Dataset Enrichment: Speaker Gender
The dataset enriches Hansard by linking parliamentary speakers to inferred gender through a multi-step external-record matching pipeline. The classification workflow then assigns stance and sexism labels using contextualized, schema-constrained LLM passes.
- Speaker gender enrichment: Speaker gender was inferred by merging EveryPolitician and MySociety records, then using honorifics, WikiData, Wikipedia, and Gender Guesser when needed.
- Speaker gender enrichment: 89.3% coverage of House of Commons speakers was achieved through exact and fuzzy matching constrained by dates and parliamentary titles.Gender Guesser achieved 97% accuracy on the study data.
- Classification workflow: Stance was classified in Pass 1, while sexism was classified in Pass 2 only for speeches with a non-irrelevant stance.Each target speech was paired with up to five preceding and five following speeches as context.
- Classification workflow: Context was used to resolve references, identify responsive arguments, and detect irony, while only the target speech determined its labels.
- Stance labels: Stance labels were for, against, both, or irrelevant, with procedural objections remaining for and grudging acknowledgement remaining against.
- Sexism labels: Hostile and benevolent sexism were treated as independent flags, each requiring a supporting target-text quotation and applicable subcategory.Hostile sexism includes degradation or control; benevolent sexism uses admired traits to justify restricting women’s roles.
Output Format
The appendix specifies schema-constrained outputs and evaluates classification robustness through cross-model agreement, sentiment checks, gender-era controls, and human-label agreement.
- Output Format: Pass 1 returns stance, rationale, and confidence; Pass 2 returns hostile and benevolent flags, subcategories, and supporting quotes.
- D Cross-model agreement: Three additional LLMs received identical prompts to assess cross-model agreement with human labels on Dvalidation.The models were GPT-5, Gemini 2.5 Flash, and DeepSeek V3.
- E Sentiment Confound Analysis: Both hostile and benevolent sexist speeches were overwhelmingly negative in sentiment across stance categories, so sentiment alone cannot distinguish them.This check used DistilBERT fine-tuned on SST-2 and all 886 sexist speeches.
- F Gender-Era Confound Analysis: Gender remained a significant predictor of supportive stance after controlling for decade (OR = 2.01, p = 0.002).
- G Sexism Classification Agreement: Table 12 reports Cohen’s κ for sexism flagging on Dvalidation, comparing inter-annotator agreement with Claude’s agreement against final human labels.
H Noise Induction Robustness Check
A 3% random gender-label noise test preserved the gender difference in stance across all 1,000 iterations. Annotation examples show that human consensus resolves ambiguous or context-dependent cases by considering both stance and gendered framing.
- Noise robustness: 3% label noise left the gender difference statistically significant (p < 0.05) in all 1,000 iterations.The robustness check randomly flipped 80 of 2,692 gendered relevant speeches in each iteration.
- Human annotation examples: Support for women police officers was classified as both supportive and sexist because the speech confined women to special duties rather than ordinary police work.The consensus identified Competitive and Complementary Gender Differentiation.
- Human annotation examples: One ambiguous speech was resolved as For and None when annotators interpreted constituency-level difficulty as compatible with support for female suffrage.Discussion gave the stance label the benefit of the doubt.
- Human annotation examples: A speech about electoral support was resolved as For and Benevolent when it praised women’s distinct qualities while supporting participation.The example frames women through complementary gender differentiation.
- Human annotation examples: Women’s participation in public roles was labelled supportive but benevolently sexist when advocacy also restricted roles or excluded married women protectively.The consensus categories were Benevolent-Protective Paternalism and Complementary Gender Differentiation.
I.2 Human-LLM (Claude Sonnet 4.6) Disagreement
Claude Sonnet 4.6 achieved agreement with human-consensus labels comparable to human annotator agreement, but disagreement cases reveal challenges in interpreting procedural, sarcastic, and appeasing rhetoric.
- Agreement: Claude Sonnet 4.6 reached Cohen’s κ = 0.711 against human-consensus labels, compared with κ = 0.644 between the two annotators.The model’s agreement was described as comparable to annotator agreement.
- Disagreement cases: A speech combining stated suffrage support with opposition to immediate action received a Both stance label from the LLM but was judged Against by humans.Humans also identified Hostile and Benevolent Paternalism, while the LLM identified Benevolent Gender Differentiation.
- Disagreement cases: Human annotators labelled a procedural suffrage speech Against and Hostile, whereas the LLM treated it as Irrelevant and None.The disagreement concerned whether separating suffrage from a bill was a delaying tactic or merely procedural reasoning.
- Disagreement cases: The LLM labelled a sarcastic proposal about Indian women’s franchise For and Benevolent, while humans classified it as Against and Hostile+Benevolent.The model interpreted positive engagement with the speaker as support and missed the hostile framing identified by humans.
J Classification Reliability
Sexism classification is difficult in long-form archaic rhetoric, but human-label checks and robustness analyses reproduce the paper’s central contrast between pro- and anti-suffrage speeches. Filtering and speaker concentration vary across the corpus, yet the reported comparisons remain broadly stable.
- Annotation reliability: Cohen’s κ below 0.4 indicates low human agreement on sexism annotation, reflecting the complexity of long-form archaic rhetoric.The authors reran the stance–sexism analysis on human-consensus labels to address this uncertainty.
- Human-label validation: 76.7% of Against-speeches contained sexism versus 37.5% of For-speeches in the 300-speech human-labelled subset.The difference was statistically significant by Fisher’s exact test (p < 0.001).
- Human-label validation: 91.7% of sexist For-speeches were benevolent-only, whereas 91.3% of sexist Against-speeches involved hostility.The same stance-by-sexism pattern appeared independently of the LLM.
- Limitations: Low LLM recall of 0.43–0.46 likely underestimates absolute sexism rates, especially benevolent sexism in For-speeches.The reported corpus rates are 21% for For-speeches and 54% for Against-speeches, while flagged cases have precision of 0.77–0.86.
- Robustness of comparisons: Against-speeches were 2.0x more sexist on human labels and 2.6x on corpus labels, indicating that relative contrasts are more stable than absolute rates.The authors argue that under-detection should not reverse the comparison unless misses are concentrated in one stance.
- Filtering across eras: Irrelevant filtering was highest before 1870 and after 1928, but retention during 1870–1928 remained stable at 42–45%.The observed pattern suggests differential filtering is unlikely to drive the temporal results.
- Speaker concentration: Speaker-weighted results remained nearly unchanged, with opponents using sexist rhetoric at 48.9% versus 22.7% among supporters.The decline of hostile sexism across eras also held when each speaker counted once.