Source-linked AI summary
Effects of Answer Format Variation on Gender Bias in Large Language Models
Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber
TL;DR
Answer format may affect how gender bias and alignment with human responses are measured in LLM evaluations, but this has not been studied in detail. The paper compares three formats across two datasets and three instruction-tuned models, finding that outcomes vary systematically and can reverse model rankings. It concludes that bias and alignment are format-contingent behaviours requiring multi-format evaluation.
Problem
The study addresses limited evidence on how answer formats affect measured LLM gender bias and alignment with human response distributions.
Method
The authors compare closed-ended, Likert-scaled and open-ended versions of questions from BBQ and OpinionQA across three instruction-tuned models.
Results
Across two datasets and three models, evaluation outcomes varied systematically with answer format, often reversing relative model rankings.
Takeaways & Limitations
Bias and alignment should be treated as format-dependent behavioural constructs, making multiple formats important for robust evaluation.
Takeaways & Limitations
The study evaluates only three small open-weight models, so its findings establish format effects without definitively estimating their magnitude across model classes.
Abstract
from arXiv · showhide
Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.
1 Introduction
The study investigates how closed-ended, Likert-scaled, and open-ended answer formats affect gender-bias measurement and LLM alignment with human survey responses. It addresses a gap in evaluations that commonly use single-format assessments despite known sensitivity to phrasing and formatting.
- Motivation: Bias metrics are highly sensitive to semantically irrelevant changes such as paraphrases, negations, answer order, and punctuation.Prior metastudies reported this sensitivity in social-bias evaluations.
- Motivation: Answer-option format remains insufficiently studied, although closed-ended and open-ended assessments can produce inconsistent and weakly correlated estimates.Existing evaluations commonly rely on single-format assessments.
- Study design: The study compares parallel closed-ended, Likert-scaled, and open-ended versions of questions from BBQ and OpinionQA across multiple instruction-tuned models.BBQ is a question-answering benchmark, while OpinionQA derives from U.S. public-opinion surveys.
- Study design: The three formats provide complementary settings: predefined categories enable direct quantitative comparison, whereas Likert scales elicit graded judgements while preserving structured comparability.The study frames answer format as part of the evaluation setting rather than merely a presentation choice.
- Contributions: The contributions are a systematic analysis of format effects, openly available multi-format datasets, and cross-format comparisons between LLM outputs and human survey responses.These contributions examine how question framing and response framing jointly relate to LLM behaviour and human-like response distributions.
2 Related Work
Related work measures gender bias through bias benchmarks and comparisons between LLM outputs and human survey distributions. It also shows that answer formats shape responses, while the consequences of changing formats for model evaluation remain insufficiently tested.
- Gender bias measurement: Bias benchmarks measure gender bias through ambiguity-sensitive fairness, stereotypical associations, and prediction invariance.BBQ operationalises fairness as appropriate uncertainty, while CrowS-Pairs, StereoSet, and Winogender probe distinct bias behaviours.
- Survey-based evaluation: Survey-based evaluations compare LLM outputs with human attitude distributions from public-opinion questionnaires.OpinionQA uses U.S. surveys, while GlobalOpinionQA extends this approach to global public opinion.
- Answer format effects: Survey research shows that question wording, option ordering, scale design, and labelling systematically affect reported attitudes and measurement validity.These effects motivate treating answer format as more than a presentational choice.
- Answer format effects: LLM evaluations often preserve original survey formats without testing whether alternatives change conclusions, despite models’ sensitivity to wording and formatting.This gap is central when datasets are used to assess model behaviour and alignment with human attitudes.
- Answer format effects: Restricting responses to simplified or binary options can constrain nuance, and no format can yet be assumed to function equivalently for models and humans.Prior work therefore treats reliance on a single prompt or format as insufficient for representing model behaviour.
3 Methodology
The study evaluates three answer formats across two gender-related datasets and three instruction-tuned models. It modifies question formats, accounts for answer-order effects, and queries each condition repeatedly to analyze response variability.
- Experimental setup: The setup evaluates all combinations of three answer formats, two datasets, and three models.The formats are closed-ended, Likert-scaled, and open-ended; the datasets are BBQ and OpinionQA.
- Models: The study uses Mistral7B, Llama8B, and Gemma12B, three instruction-tuned models in the 7–12B parameter range.Instruction-tuned models are selected for responding within predefined formats while providing architectural and training diversity.
- Datasets: The data comprise 360 gender-related BBQ questions and 158 manually extracted gender-related OpinionQA items.The BBQ sample balances ambiguous and disambiguated contexts, while OpinionQA human distributions serve as the empirical reference.
- Format modification: Each question is adapted into closed-ended, Likert-scaled, and open-ended variants with format-specific constraints and response requirements.Closed-ended and Likert-scaled variants retain answer-set constraints, whereas open-ended variants remove answer options and may manipulate reasoning instructions.
- Format modification: All closed-ended variants permute answer options to account for position bias, with six permutations tested per OpinionQA item.BBQ uses all six permutations, producing 2,160 total variants; OpinionQA uses cyclic rotations and reversed orderings.
- Response generation: Each model is queried ten times per question and format condition at temperature 0.7 to capture response variability.Open-ended responses are mapped to the original multiple-choice categories using Qwen2.5-32B-Instruct and checked through manual inspection of random samples.
4 RQ1: Gender Bias Measurement
Gender-bias measurements vary substantially with answer format because closed-ended, Likert-scaled, and open-ended responses elicit different behavioural patterns. No single format provides a complete or stable characterisation of model behaviour, and minor design choices can alter model comparisons.
- Cross-format synthesis: Mistral7B shifts from the most biased model in closed-ended evaluation to the least biased in open-ended evaluation, demonstrating cross-format reversals in model rankings.Gemma12B has near-zero directional net bias for disambiguated questions but increased stereotype-consistent bias for ambiguous questions in both closed-ended and open-ended formats.
- Closed-ended: Closed-ended formats show the strongest and most consistent stereotypical behaviour across models, with effects differing between identity-label and proper-name subsets.Mistral7B and Llama8B show significant effects across both subsets, while Gemma12B has near-zero sDIS but substantial sAMB, especially for proper names.
- Closed-ended: Answer-order permutations meaningfully affect accuracy and measured bias in closed-ended evaluation, sometimes reversing the apparent direction of sDIS.Binary versus multi-point scales and other seemingly minor design choices also produce material differences in model comparisons.
- Likert-scaled: Likert formats can produce near-zero sAMB despite substantial polarisation, as stereotype-countering and stereotype-reinforcing responses partially cancel in the mean.Across all three models, ambiguous responses frequently span both extremes, while low odd-scale accuracy suggests limited midpoint use as an escape option.
- Open-ended: Open-ended formats substantially reduce sAMB and consistently increase ‘UNKNOWN’ rates, indicating more frequent refusal or non-substantive answering.The ‘UNKNOWN’ category combines ambiguity recognition, explicit refusals, non-commitment, and rare off-target responses, so lower measured bias may not indicate lower underlying bias.
5 RQ2: Human Opinion Alignment
OpinionQA measures alignment as distributional similarity between LLM and human responses, using metrics that differ in their treatment of ordinal and categorical structure. Answer format substantially changes measured alignment and model rankings, while metric and preprocessing choices also shape the scores.
- Alignment measurement: OpinionQA operationalises human alignment as distributional similarity between model outputs and human survey responses for equivalent gender-related questions.The evaluation uses Wasserstein distance, Jensen-Shannon distance, and composite Dα to capture ordinal and categorical similarity.
- Alignment measurement: Likert responses are transformed into normalised distributions, while open-ended responses are annotated and scored like closed-ended responses.Likert ratings are reversed for negative framing, averaged across samples, and renormalised to sum to 1 before comparison with human responses.
- Cross-format synthesis: Answer format produces substantial alignment shifts, including complete reversals in model rankings across formats.Likert values are closer to the human reference distribution, but Mistral7B aligns best on discrete formats while Gemma12B shows the reverse pattern.
- Cross-format synthesis: Structural format properties mainly alter ordinal probability-mass placement, whereas broader model-specific distributional tendencies remain comparatively stable across formats.Answer order and scale granularity are identified as the main structural influences in the Wasserstein and JS decomposition.
- Metric interpretation: Alignment metrics are not format-invariant because scores reflect interactions between model outputs and measurement assumptions.Lower Likert divergence partly reflects mean-normalisation and other mechanical factors, while JS distance is sensitive to sparse categories.
- Format-specific findings: Closed-ended outputs show systematic answer-order effects, Likert outputs show pervasive scale-condition effects, and open-ended composite distances fall between the two formats.Binary Likert scales differ substantially from other scales, while JS contributes roughly two-thirds of the open-ended composite metric.
6 Conclusion
Answer format systematically shapes measured gender bias and human-distribution alignment in LLMs, including reversals in model rankings. The findings frame bias and alignment as format-dependent behaviours, requiring context-sensitive, multi-format evaluation.
- Conclusion: Across two datasets and three instruction-tuned models, evaluation outcomes varied systematically by answer format and often reversed relative model rankings.The study examined closed-ended, Likert-scaled and open-ended formats for gender-bias measurement and alignment with human opinion distributions.
- Conclusion: Open-ended formats produced weaker bias signals, while model-based annotation exposed a trade-off between ecological validity and measurement reliability.Formats that better capture everyday usage may be less amenable to systematic quantitative analysis.
- Conclusion: Repeated sampling reflects stochastic variation in model outputs rather than between-individual heterogeneity, limiting direct comparison with human response variance.This difference complicates interpreting repeated model responses as equivalent to surveying multiple human respondents.
- Conclusion: Bias and alignment are format-dependent behavioural constructs, with formats capturing different aspects including directional preferences, polarisation and abstention.Single-format evaluations therefore capture only part of the underlying multidimensional construct without reducing gender bias to a measurement artefact.
- Conclusion: Bias and alignment emerge from interactions between model representations and response constraints, so the relevant format depends on the evaluative or deployment context.The methodological question is not which format is universally most accurate, but which behavioural manifestation matters for the intended context.
Limitations
The study’s conclusions are limited by its narrow model, bias, gender-category, language, and cultural scope, as well as methodological and evaluation constraints. These limitations complicate generalisation and the interpretation and comparability of measured bias and distributional alignment.
- Scope: The study evaluates only three small open-weight models and one bias dimension, so its findings are evidence that format effects exist rather than definitive estimates across model classes or bias dimensions.Generalisability to larger models and dimensions such as race, disability status, and intersectionality remains open.
- Scope: Binary male/female categories capture only a limited subset of possible gender-related biases.This choice follows OpinionQA’s survey structure and was applied consistently to BBQ.
- Scope: English prompts grounded in a U.S. socio-cultural context may limit applicability elsewhere, where bias expression and response-scale interpretation may differ substantially.The authors encourage broader coverage across linguistic and cultural settings and other dimensions.
- Methodological limitations: Changing answer format also changes framing and decision processes, while Wasserstein and JS distance rely on assumptions that may impose artificial structure or overestimate differences.Likert scales remove mutual exclusivity; Wasserstein assumes ordinal structure, whereas JS distance is sensitive to sparsity and rare categories.
- Evaluation limitations: Subset analyses limit direct comparison with prior studies, and the functional ‘UNKNOWN’ approach plus LLM label-based annotation can obscure or misclassify bias signals.Scores are not directly comparable estimates of overall representativeness; non-neutral Likert responses may conflate counter-stereotypical bias with uncertainty, while open-ended annotation may miss indirect expressions.
Ethical Considerations
The study raised no external approval requirements because it involved no human participants and used publicly available datasets and open-weight models.
- Ethical Considerations: No human participants were involved, and the publicly available BBQ and OpinionQA datasets and open-weight models required no external ethical approval.BBQ is licensed CC-BY 4.0, OpinionQA permits statistical and scientific research use, and all models were accessed via Hugging Face.
A Dataset Construction · B Generation of Likert Scales · C Response Generation and Collection
The study constructs adapted BBQ and OpinionQA datasets, generates Likert versions with model-assisted prompts and validation, and collects closed-, Likert-, and open-ended responses under format-specific procedures. These procedures standardize decoding where possible while accommodating model output variation, annotation needs, and dataset-specific answer structures.
- A Dataset Construction: BBQ construction retained within-group gender comparisons, selected one manually frequency-informed name pair per question index, and sampled uncertainty options during dataset construction.Intersectional pairings were removed, and the resulting subset was not uniformly balanced across ten expressions.
- A Dataset Construction: OpinionQA preprocessing minimally corrected 19 questions, redacted 8 survey-specific items, reduced regional specificity in 20 questions, and reworded 14 WHYNOTBIZF2*_W36 questions.The edits aimed to preserve core meaning while making prompts more concise and less regionally specific.
- B Generation of Likert Scales: BBQ Likert scales began with a 4-point negative scale generated by Qwen2.5-32B-Instruct at temperature 0, after which remaining scales were constructed programmatically.Generation used an NVIDIA RTX 6000 Ada GPU (48 GB).
- B Generation of Likert Scales: OpinionQA Likert generation used multiple demonstrations, guidance for complete answer statements, validation, and manual correction of grammatical issues.Questions with WHYNOT*_W36 keys received a separate generation round, explicit structural guidance, and additional manual editing.
- C Response Generation and Collection: All response-collection models used temperature 0.7 and top-p 0.9, while Qwen2.5-32B-Instruct also performed open-ended response annotation.Inference ran on an NVIDIA RTX 6000 Ada GPU (48 GB), and stochastic decoding was used because multiple responses could be plausible in underspecified settings.
- C Response Generation and Collection: Closed-ended BBQ and OpinionQA collection instructed models to choose exactly one letter, with OpinionQA split into rounds according to answer-set length.The letter-based format remained constant, while prompts were adjusted to match the number of OpinionQA options.
- C Response Generation and Collection: Likert collection mirrored closed-ended querying but required exactly one number from the provided scale, retaining MAX_NEW_TOKENS=16 for output stability.Mistral7B frequently ignored brevity instructions, although no manual postprocessing was needed.
- C Response Generation and Collection: Open-ended BBQ responses used MAX_NEW_TOKENS=256 and reasoning/no-reasoning variants before Qwen2.5-32B-Instruct mapped outputs to original options; OpinionQA used MAX_NEW_TOKENS=512 and fuzzy matching fallback.Mistral7B annotation failures in OpinionQA were logged rather than handled identically to other outputs.
D BBQ Metrics
BBQ bias metrics quantify directional stereotype reinforcement while distinguishing disambiguated from ambiguous contexts. Scores use normalized response polarity and are reported across item subsets to capture differences between abstract labels and named individuals.
- Disambiguated contexts: Disambiguated bias scores range from −1 for consistently stereotype-countering responses to +1 for consistently stereotype-reinforcing responses, with 0 indicating no net bias.The score is calculated among non-‘UNKNOWN’ outputs, where biased answers align with social bias.
- Ambiguous contexts: Ambiguous bias scales directional bias in erroneous non-‘UNKNOWN’ answers by the error rate, reflecting that substantive answers are incorrect in ambiguous contexts.A model that always selects ‘UNKNOWN’ receives sAMB = 0 regardless of latent bias.
- Response normalization: Responses are normalized to a common polarity scale b(ri) ∈[−1, +1] measuring the direction and magnitude of stereotype reinforcement.Closed-ended and open-ended group choices have magnitude ±1, ‘UNKNOWN’ has magnitude 0, and Likert responses use normalized agreement strength.
- Reporting subsets: sDIS, sAMB and accuracy are reported for the full sample, gender-identity labels, proper names and aggregate results across both subsets.This distinguishes abstract group labels from named individuals when examining bias patterns.
E BBQ Granular Results
Across BBQ answer formats, measured bias and accuracy vary substantially: response ordering alters closed-ended conclusions, Likert scales reveal polarized stereotyping despite near-zero mean bias, and open-ended responses reduce measured bias while changing accuracy. These formats therefore elicit distinct response patterns rather than interchangeable measurements.
- Closed-ended: Gemma12B combines near-zero sDIS with significant sAMB, while Llama8B and Mistral7B show stereotype-reinforcing bias alongside lower accuracy.Gemma12B produces 79–84% ‘UNKNOWN’ responses; Llama8B and Mistral7B have approximately 50% and 64% accuracy, respectively.
- Closed-ended: Closed-ended answer ordering substantially changes both BBQ accuracy and apparent bias, with model-specific rankings and reversals across permutations.Llama8B and Mistral7B perform best when ‘UNKNOWN’ appears first, whereas Gemma12B performs best when it appears last.
- Likert-scaled: Likert responses show near-zero sAMB but substantial positive polarisation, as stereotype-reinforcing and counter-stereotypical extremes largely cancel in the mean.Gemma12B’s combined sAMB is −0.1%, with polarisation of 24.4% at τ = 0.6; identity-label and proper-name polarisation are 27% and 21.8%.
- Likert-scaled: Likert scale length and polarity produce few significant pairwise differences in sDIS or sAMB, primarily shaping response distributions rather than overall bias scores.Llama8B and Mistral7B nevertheless show significant net stereotype-reinforcing bias, with polarisation indices of 22.3% and 34.7%.
- Open-ended: Open-ended responses yield non-significant sDIS and substantially reduced sAMB, while accuracy changes markedly across models: 66.3%, 80.8%, and 97.5%.These accuracy values correspond to Gemma12B, Llama8B, and Mistral7B, respectively, compared with their closed-ended results of 79%, 50.1%, and near-ceiling performance.
- Open-ended: Most ‘UNKNOWN’ annotations in open-ended responses reflected genuine uncertainty or insufficient information, although occasional errors affected a small minority of items.The open-ended BBQ condition mapped expressions of uncertainty to ‘UNKNOWN’, unlike the OpinionQA condition where explicit refusals were distinct.
F OpinionQA Metrics
OpinionQA compares model and human response distributions using Wasserstein distance, which incorporates ordinal structure, and Jensen–Shannon distance, which treats responses categorically. Because the dataset’s ordinal encodings can impose questionable orderings on qualitatively distinct answers, Wasserstein results require caution, motivating a complementary categorical metric and composite measure.
- Wasserstein distance: Wasserstein distance measures distributional differences using the ordinal structure of OpinionQA response categories.Each substantive answer option receives a dataset-inherited ordinal value, and the normalised distance lies in [0, 1].
- Jensen–Shannon distance: Jensen–Shannon distance treats responses as categorical distributions, includes ‘Refused’, and avoids assumptions about category ordering.It is computed in base 2, yielding values in [0, 1], with ε added before renormalisation to handle zero probabilities.
- Wasserstein distance: The Wasserstein encoding assigns neutral values of 1.5 for 84 of 134 four-option questions and 2.5 for both six-option questions.These values are pre-specified mappings inherited from the original OpinionQA dataset.
- Wasserstein distance: Wasserstein distance can impose an ordering on qualitatively distinct choices, so its results should be interpreted cautiously.For example, options concerning having children early, waiting until career establishment, or not having children are encoded as {1.0, 2.0, 3.0}.
- Composite distance: The composite distance uses α = 0.5 to weight Wasserstein and Jensen–Shannon distances equally, balancing ordinal sensitivity with general distributional similarity.Because both component distances lie in [0, 1], the composite metric also lies in [0, 1].
G OpinionQA Granular Results
OpinionQA results show that answer format substantially changes distributional alignment: closed-ended responses exhibit ordering effects, Likert responses exhibit strong scale-length effects, and open-ended responses fall between the other formats. Reasoning effects are smaller than format effects overall, though they improve some models and reduce refused or off-target responses.
- Closed-ended: Closed-ended responses show systematic order effects across models in the 4-option subset, with Gemma12B varying modestly and Mistral7B reaching its lowest scores under the 2341 ordering.For Gemma12B, WDist ranges from 0.267-0.285 and JSDist from 0.547-0.568; the 2341 ordering yields the lowest WDist.
- Closed-ended: Smaller subsets cannot reliably separate ordering effects from question-content effects, while degenerate confidence intervals suggest deterministic response mappings in several low-sample orderings.The 6-option subset has n = 2, and several 3-option orderings have n = 13 for Gemma12B and Mistral7B.
- Likert-scaled: Likert alignment deteriorates substantially with 2-point scales, whereas 4-point and longer scales remain comparatively stable, indicating that precision saturates at four points.Scale length matters more than polarity, with within-scale polarity comparisons generally non-significant.
- Likert-scaled: Model-specific Likert patterns include Gemma12B’s agreement skew, Llama8B’s highest alignment at 4 points and |∆|composite = 0.137 polarity gap, and Mistral7B’s non-significant negative-polarity trend.The polarity gap for Llama8B occurs at 2 points and is absent in the other models.
- Open-ended: Open-ended composite distances lie between closed-ended and Likert formats, with JSDist dominating because annotations produce discrete distributions.Reasoning significantly improves Gemma12B and Mistral7B by |∆|composite = 0.025 and 0.036, respectively, mainly through JSDist; Llama8B changes non-significantly in the opposite direction.
- Open-ended: Reasoning consistently reduces refused or off-target responses across models, while excluded non-compliant responses remain rare at 30 instances total.Observed non-compliant cases include one Llama8B refusal and Mistral7B off-scale values in 2-point conditions; classified ‘Refused’ responses were retained as meaningful observations.