Source-linked AI summary

The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, Scott A. Hale

arXiv:2404.16019v2cs.CL

TL;DR

Existing human-feedback datasets provide limited demographic coverage, fine-grained explanation, and contextual information about annotators. PRISM links surveys and live conversations for 1,500 participants across 75 countries, then uses the data to show that alignment findings vary with sample composition and discussion topics. The dataset therefore supports more individualised and multicultural analysis, while retaining privacy, harmful-content, and sampling limitations.

  • Problem

    Human-feedback datasets often rely on binary comparisons, narrow samples, limited participant information, and weak documentation of how values are operationalised.

  • Method

    PRISM maps survey-based participant profiles and stated preferences to fine-grained ratings from live conversations with LLMs, including diverse recruitment, UK and US census-representative samples, and pseudonymous individualised records.

  • Results

    Rankings vary with sampled participants and conversation topics, and PRISM rankings differ from other leaderboards because of dataset and response characteristics.

  • Takeaways & Limitations

    Alignment data should be interpreted with careful attention to which people provide it, what they discuss, and how their preferences are aggregated.

  • Takeaways & Limitations

    PRISM retains privacy risks and may yield different outcomes under alternative data divisions, while controversy-focused collection can encourage harmful content.

Abstract

from arXiv · show

Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a dataset that maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide what alignment data.

1 Introduction

PRISM addresses gaps in human-feedback datasets by linking diverse participants’ characteristics and preferences to contextual, fine-grained feedback on LLM conversations. It is designed to study subjective and multicultural alignment without reducing values to generic human preferences.

  • Existing datasets often rely on binary comparisons, narrow samples, limited participant information, and poorly documented value operationalisation.
  • The dataset includes participants from 75 birth countries, 8,011 conversations, 21 models, 27,172 interactions, and 68,371 utterances.
  • Generic human data can produce reductionist behaviours because values depend on people, communities, and operating contexts.
  • PRISM maps survey responses from 1,500 diverse participants onto live LLM conversations, supporting both contextual-preference comparisons and stated-preference approaches.
  • PRISM combines participatory recruitment, UK and US census-representative samples, and individualised ratings linked to pseudonymous participant profiles.

2 The PRISM Alignment Dataset

PRISM uses a two-stage design: participants first provide profiles and stated preferences, then converse with multiple LLMs and rate responses using contextual and fine-grained feedback. The protocol deliberately varies conversation objectives while preserving participant choice and documenting measurement trade-offs.

  • Stage 1: Survey: Participants complete a survey covering demographics, LLM familiarity, stated behavioural preferences, and personal values before entering live conversations.
  • Stage 2: Conversations: Conversation prompts are collected under unguided, values-guided, and controversy-guided conditions to diversify topics across the subjective spectrum.
  • Stage 2: Conversations: Participants receive up to four model responses to an opening prompt, rate them from 1–100, and provide both performance and choice-attribute feedback.
  • Design boundaries: The released data controls for deviations from the intended two-conversations-per-condition quota, while failed or slow model responses may leave participants with fewer than four opening responses.
  • Measurement choices: Cardinal ratings express preference intensity and can be converted to rankings, but they introduce intrapersonal noise and reduce interpersonal comparability.
  • Stage 2: Conversations: The highest-scoring model is locked into subsequent turns, where participants continue for 2–10 turns and rate two nondeterministic responses from that model.

3 Experiments with PRISM

PRISM’s experiments show that both discussion topics and model rankings vary with participant identities, conversational context, and sampling choices. Representative, larger samples generally produce better welfare outcomes, while narrow samples can disadvantage out-groups and no model satisfies a majority.

  • Case Study I: 70% of prompts fall into 22 topic clusters, while 30% remain outliers after embedding, dimensionality reduction, and HDBScan clustering.Identity-group topic prevalence is then estimated with OLS regressions using clustered standard errors.
  • Case Study I: 11% of identity-related coefficients are significant at α = 99%, with topic differences linked to gender, age, race, and region.Local prompt neighbourhoods are intersectionally diverse: 84% meet or exceed entropy expected under random sampling.
  • Case Study II: Rankings vary with participant samples, conversational topics, and geographic groups: palm-2 drops 4 places in the US, llama-7b drops 7 in Asia, and mistral-7b gains 7 in Africa.The analysis uses 6,669 balanced conversations and Pairwise Rank Centrality to aggregate model preferences.
  • Case Study III: Smaller samples increase the probability of selecting a model with worse mean welfare, while exclusive sampling from a group tends to reduce out-group welfare.The welfare analysis compares representative samples of 10, 20, 50, or 100 people with demographically restricted samples of 100.
  • Case Study III: 45% is the maximum MEANCHOICE probability achieved by a model in the US, so no single model reaches majority preference.If the winning model is shown alongside three randomly selected alternatives, participants choose it with probability below 50%.

4 Related Work

Related work frames human-feedback alignment as a problem involving feedback signals, training methods, and participation. It highlights the risks of narrow participation and collapsing subjective variation into generic human preferences.

  • Participation & Representation in Science & Technology: Collapsing subjective experience into majority votes can hide annotator disagreement, while over-generalising from WEIRD societies risks treating culturally specific conclusions as universal.PRISM releases participant IDs and characteristics to expose sample diversity and acknowledge sample specificity.
  • Learning from Human Feedback: Human-feedback pipelines use comparisons, principles, fine-grained feedback, or natural language to reward dimensions such as helpfulness, honesty, and harmlessness.Reward models can update LLMs with PPO or Reinforce, while DPO and supervised fine-tuning provide reward-model-free alternatives.

5 Limitations, Discussions and Conclusions

PRISM broadens human-feedback data through participatory, individualised collection while exposing unresolved ethical and methodological boundaries. Its diversity and openness also introduce privacy, harm, variance, statistical-power, and representativeness concerns.

  • Ethical boundaries: Privacy risks remain because PRISM contains sensitive conversations, despite informed consent, pseudonymisation, PII checks, and anti-deanonymisation terms.The dataset includes conversations involving topics such as abortion, religion, immigration, workplace disputes, and intimate relationships.
  • Ethical boundaries: Controversy-focused conversations expand preference data into disagreement-rich domains but may encourage hateful, bigoted, biased, or otherwise harmful content.The authors report that PRISM is less toxic than previous datasets while retaining concerns about harmful content.
  • Methodological boundaries: Free dialogue, cardinal feedback, and varied models and participants increase diversity and subjective freedom but complicate controlled experiments and limit statistical power.The authors also note that alternative divisions of the data may produce different outcomes.
  • Representativeness: PRISM remains biased toward English-speaking crowdworkers whose task-specific incentives may not align with wider populations.This scope boundary limits how broadly the dataset’s participants can represent human populations.
  • Discussion: Preference-based utilitarianism may diverge from individual or societal well-being, making disagreements and participant positionality relevant to aggregation decisions.The discussion highlights conflicts between personal preferences, self-interest, and others’ interests.
  • Contributions: PRISM contributes a dataset from 1,500 diverse humans to inclusive, participatory research on aligning LLMs with human preferences in a pluralistic world.The authors frame alignment data as a public good and position PRISM as an open resource for this research question.

Author Contribution Statement

The contribution statement assigns distinct roles across conception, data collection, development, analysis, writing, editing, annotation, metadata processing, and codebase work. Manuscript editing and feedback involved everyone.

  • Roles: Project conception was credited to Kirk, Hale, and Vidgen.
  • Roles: Data collection design was credited to Kirk, Hale, Vidgen, Röttger, and Margatina.
  • Roles: Frontend and backend development were credited to Kirk and Ciro, and Kirk and Mosquera, respectively.
  • Roles: Analysis advisory was credited to Hale, Vidgen, Röttger, Bartolo, Bean, Williams, and He.
  • Roles: Literature and dataset comparison, metadata processing, manual annotation, and results and codebase work were assigned to the listed contributors.
  • Roles: Manuscript writing was credited to Kirk and Whitefield, while manuscript editing and feedback involved everyone.

Supplementary Material

The supplementary materials document PRISM’s access points, file formats, metadata, participant and model characteristics, and ethical safeguards. They also expose important boundaries around language, sampling, detection, privacy, and harmful content.

  • Data access: PRISM is available through GitHub and HuggingFace, with permanent DOI 10.57967/hf/2113.
  • Data organisation: The dataset includes survey, conversation, utterance, and metadata JSON Lines files, supported by codebooks and format-conversion code.Survey rows identify participants, conversation rows identify conversation trees, and utterance rows represent scored interactions.
  • Data organisation: Metadata records language detection, PII detection, and moderation flags separately from the main data files.
  • Participants: All participants were recruited through Prolific for academic research on how people interact with LLMs and perceive their outputs.
  • Language: Although text language was not explicitly restricted, English screening and English instructions resulted in 99% of text instances being in English.Participants were born in 75 countries and resided in 38 countries, but represented English varieties were not documented.
  • Sampling: PRISM improves diversity relative to early feedback datasets but still skews White, Educated, and Western because all participants came from one crowdworking platform.The platform introduces biases associated with internet use, hourly payment, and self-selection.
  • Models: The dataset contains 21 models across families, capabilities, and sizes, including 12 commercial APIs and 9 open-access models.
  • Detection: The AI-text detector produced potentially unreliable classifications, including an 88.1% score for feedback human annotators judged strongly human-generated.

B.10 Expanded Technical and Task Design Limitations

PRISM’s rich, naturalistic design creates interpretive and experimental limits. Identity categories, conversational choices, model randomness, cardinal ratings, crowdworker sampling, and model turnover all constrain generalisation and causal inference.

  • Intersectionality: PRISM’s many structured and unstructured participant attributes permit numerous data divisions, but sparse combinations are under-powered and alternative choices may yield different outcomes.Aggregating groups can also lump distinct geographies together as cultures, while some regional categories remain highly concentrated.
  • Experimental control: Participant-selected topics, randomly selected models, and stochastic model responses create many moving sources of variance that hinder identifying robust mechanisms of preference differences.Free choice was retained to model naturalistic LLM use and dialogue diversity.
  • Measurement: Subjective cardinal ratings may provide noisy interpersonal signals because measurement invariance and preference falsification can confound conclusions.The authors caution against relying too heavily on cardinal ratings for interpersonal comparisons.
  • Representativeness: PRISM still consists only of crowdworkers, limiting representativeness because participants are digitally engaged, platform-recruited, and subject to sample biases.The authors therefore avoid overstating the dataset’s diversity claims.
  • Model coverage: The pace of new model releases creates incompatibility between current model coverage and the lengthy ethics, interface, processing, and annotation processes required for participant research.

C PRISM Data Clause

PRISM is provided as a research and educational dataset with explicit usage, safety, attribution, privacy, and maintenance conditions. Its collection involved consented participant surveys and LLM conversations, with moderation checks and documented annotation limits.

  • PRISM is intended for research and educational use in natural language processing, conversational agents, social science, and related AI applications.
  • Users must follow model-specific terms, apply filtering and moderation to potentially unsafe or offensive conversations, and avoid deanonymising individuals or entities.
  • Human-written texts use CC-BY-4.0, while model responses use CC-BY-NC-4.0 and remain subject to original provider licenses.
  • Participants provided informed consent and completed a survey followed by conversations in which they prompted and rated AI language models.
  • The dataset includes automated moderation metadata, but manual review found no actual PII in 167 participant-written texts flagged for PII.
  • Ethnicity and religion annotations involve unavoidable subjectivity, so future analyses should choose categorisations according to the research question.
  • PRISM recruits English-fluent participants born and residing in the same country, while retaining participants whose reported birth and residence countries differ.

L Census Rebalancing

PRISM rebalances UK and US samples toward census proportions, improving but not eliminating demographic mismatch. Rebalancing reduces sample sizes and leaves important representation gaps and unmeasured characteristics uncontrolled.

  • Cyberattacks disrupted some workflows, causing timed-out participant spots to be refilled by others with the same demographics.
  • 300 is the target rebalanced sample size, but the final samples contain 243 UK and 230 US participants because available data constrain resampling.
  • After rebalancing, demographic differences in both samples fall within approximately 7 percentage points, but the UK sample still lacks older participants and the US remains demographically skewed.
  • Rebalancing uses UK 2021 and US 2022 census data for age, ethnicity, and gender comparisons, with US “Other” combined with Hispanic.
  • Improving representativeness on observed census characteristics reduces sample size and may worsen representation of unobserved characteristics such as political affiliation, education, or income.

M Text and N-Gram Analysis

PRISM combines structured and free-text analyses of participant preferences, feedback, prompts, and model responses. These analyses characterize language patterns, score usage, conversation lengths, and additional behavioural priorities.

  • Text distributions: PRISM summarizes five free-text instance types with counts, length distributions, frequent N-grams, and adjective patterns for selected text fields.The analysis covers constitutions, self-written profiles, open feedback, prompts, and model responses.
  • N-gram analysis: User prompts emphasize information-seeking dialogue and questions, while model responses show an advisory tone and frequent de-anthropomorphisation.These patterns are reported from the top N-grams in prompts and responses.
  • Preference relationships: Participants’ stated preferences are not highly correlated with contextual choice or performance attributes.The paper notes this may reflect difficulty specifying preferences or differences between stated and observed behaviour.
  • Additional behavioural attributes: Open feedback identifies user adaptation, cultural adaptation, bias correction, factuality, limitation awareness, temporal updates, anthropomorphic presentation, and accessibility as behavioural priorities.Participants also expressed contrasting preferences about human-like behaviour versus explicit disclosure that the system is AI.
  • Scores and conversation length: Most conversations have two turns, while longer conversations involve fewer participants; after the first turn, mean scores increase and score ranges fall.Raw scores also show interface-related bunching at 1, 50, and 100, and continuers use a narrower, positively skewed range.

R.2 Topic Prevalence Regression Results

Topic prevalence varies with conversation type and, to a lesser extent, participant demographics. However, most regression coefficients are nonsignificant, and the specification explains only a small share of variation in topic choice.

  • Regression results: 16% of 682 tested coefficients are significant at α = 99%, falling to 11.4% for demographic affiliations after controlling for conversation type.The controlled analysis contains 565 nonsignificant and 73 significant demographic relationships.
  • Demographic associations: Women and non-binary participants discuss gender and LGBTQ+ identity more than men, while older participants discuss elections and travel more than younger participants.Older participants are less likely to discuss managing relationships or job search; Black participants discuss climate change less than White participants.
  • Regional associations: Almost all regions question LLMs about abortion less often than US participants, while regional, ethnic, and religious affiliations are partly collinear.The paper notes that the Middle Eastern association with Israel–Palestine discussion may reflect national, ethnic, or religious affiliations.
  • Model fit: Across 22 topic regressions, R2 ranges from 0.008 for Exploring AI and Machine Learning to 0.11 for Managing Relationships, with a mean of 0.03.Thus, a large proportion of topic-choice variation remains unexplained by the specification.
  • Topic clusters: The topic-clustering analysis organizes prompts into named clusters while retaining a substantial outlier set for prompts not assigned to a cluster.The supplied table caption reports that HDBSCAN leaves 32% of prompts unclustered.

S.1 Extended Methods

The method extracts semantically constrained local neighbourhoods from prompt embeddings, removes singleton and same-author groups, and measures their intersectional demographic entropy against random expectations. At τcos = 0.125, the resulting neighbourhoods are generally diverse despite their semantic similarity.

  • Local neighbourhood extraction: Single-link hierarchical clustering merges prompt embeddings within cosine distance threshold τcos, allowing neighbourhood size k to vary while constraining semantic similarity.The procedure computes pairwise cosine distances, merges qualifying neighbourhoods, consolidates IDs, and returns neighbourhood assignments.
  • Local neighbourhood extraction: Singleton neighbourhoods and non-singleton groups containing prompts from only one participant are removed before demographic analysis.The analysis is repeated for τcos values of 0.05, 0.125, and 0.2.
  • Intersectional entropy: Intersectional entropy combines adjusted per-attribute entropies across demographic attributes after accounting for group counts, population imbalance, and neighbourhood size.Expected entropy is simulated by randomly sampling k-sized neighbourhoods using population-wide probabilities.
  • Findings: 84% of neighbourhoods fall above or within the expected entropy range for an equivalently sized random sample at τcos = 0.125.Participant ID and Cluster ID provide robustness checks for non-duplicated authorship and containment within one topic cluster.
  • Findings: Only 273 of 8,011 prompts form unique local neighbourhoods, while the neighbourhoods that emerge generally contain authors diverse across geography, religion, age, and intersecting attributes.Less than 1% of prompts have no intersectional diversity, and 58% represent at least two subgroups for all five attributes.

S.3 Local Neighbourhood Robustness Checks

Robustness checks show that local neighbourhoods remain interpretable across cosine thresholds, while fixed or near-fixed dialogue contexts reveal substantial variation in participant ratings. The findings support a tie threshold of 5–10, but model-rank midsections remain sensitive to analysis choices.

  • Fixed dialogue contexts: 36.3 ± 26.5 is the score distribution across different participants rating semantically constrained contexts at τcos = 0.05.The authors caution that participants self-select into these duplicate groups, so allocation is non-random.
  • Threshold robustness: 154 neighbourhoods at τcos = 0.05 are above or within the 99% confidence interval for expected entropy, including multicultural groups discussing religion and abortion.The largest examples contain 14–60 prompts and vary only in formatting, capitalisation, or punctuation in some cases.
  • Fixed dialogue contexts: Scores of 67 and 90 for two very similar responses illustrate participant disagreement within a strict field site.The responses address the same religion question and differ only slightly in wording.
  • Fixed dialogue contexts: At least two participants rate the same prompt-response pair in 40 field sites, showing that PRISM can retrieve empirically fixed dialogue contexts.The field-site table defines score range as the maximum minus minimum after combining unique participants’ scores.
  • Rating calibration: A tie threshold of 5–10 is recommended because same-participant duplicate ratings help calibrate noise in the visual analogue scale.The recommendation is presented as sensible rather than universally required.

T.1 Extended Methods

The extended methods examine how participant, topic, scoring, aggregation, and sampling choices affect model comparisons. They combine alternative preference-processing procedures with robustness analyses of ranks, score correlates, and sample size.

  • Score processing: Participants’ score-processing choices reflect an unresolved distinction between measurement-scale differences and genuine community-level preference differences.Normalising scores can correct participant fixed effects but may also flatten signals associated with divergent communities.
  • Score processing: Tie thresholds reduce visual-analogue-scale noise but introduce an arbitrary mixture of cardinal and ordinal comparison.The analysis therefore tests rank sensitivity to the chosen threshold.
  • Preference aggregation: Preference aggregation functions differ in their assumptions about whether ratings are cardinally measurable and whether score units are comparable across participants.The paper compares functions including Elo, average win rate, mean score, and within-participant normalisations.
  • Preference aggregation: Pairwise Rank Centrality converts model ratings into win-loss comparisons and represents models as nodes in a transition graph.Ties within threshold t = 5 contribute both win-loss and loss-win comparisons.
  • Topic and demographic effects: Topic-model comparisons report male–female differences in mean model score, but many demographic cells are small and require caution.The analysis displays the difference for each topic-model pair and identifies binary gender as the largest demographic division.
  • Sample-size robustness: Increasing the participating cohort improves stability in model rank, whereas very small samples allow almost any model to rank highly.The analysis uses 1,000 bootstrap samples and includes 1,246 participants in the balanced subset.
  • Text correlates: 9% of conversations contain at least one model response matching refusal phrases, and a non-refuser is chosen 73% of the time in those cases.The analysis operationalises refusal using matched refusal phrases such as “I’m sorry, but...”.

U.2 UK Sample

The UK welfare analysis compares welfare distributions across subpopulations and sampling schemes, with results qualitatively similar to the US analysis. It also documents participant familiarity, use cases, and stated preferences collected for the sample.

  • UK welfare analysis: The UK welfare distributions use MEANWELFARE with MAXRATING and MEANCHOICE with MAXCHOICE, comparing sampling schemes against Rep (100).The figure distinguishes Rating and Choice comparisons and marks distributions that are FOSD by Rep (100).
  • UK welfare analysis: The UK welfare results are qualitatively similar to the US results in Fig. 5.
  • Imputing missing individual welfare: Missing individual welfare values are imputed with multivariate imputation using only the individual-welfare matrix for each model.The implementation uses Python’s IterativeImputer, in an approach described as similar in spirit to collaborative filtering.
  • Participant and conversation metadata: 1,500 participants are represented in the linked metadata, with pseudonymized user identifiers connecting survey and conversation data.The user identifier is pseudonymized from the Prolific worker ID, while conversation identifiers link metadata to the main data.
  • Participant survey: Participants reported varied familiarity with language models and frequency of use, while the survey collected use cases and stated preferences over model behaviours.Familiarity responses included 920 somewhat familiar, 424 very familiar, and 156 not familiar at all; use frequency was recorded for participants who had used language models.

N Missing: 0 N Unique: 1475

The dataset records participants’ values, desired model traits, demographics, affiliations, locations, and study metadata. It combines self-described responses with categorized variables and records substantial variation in participant characteristics.

  • Self-described affiliations: 264 unique ethnicity strings and 137 unique religion strings were collected through self-description, then categorized for analysis.Two independent annotators manually verified the automated classifications of self-described ethnicity and religion.
  • Categorized affiliations: The categorized sample includes 969 White, 122 Black or African, 121 Hispanic or Latino, and 95 Asian participants, while religion categories include 762 non-religious and 487 Christian participants.
  • Geographic metadata: Participants came from 75 countries of birth and 38 countries of residence, with adjusted regional categories separating the UK and US from broader regions.The adjusted residence categories include 338 US, 313 Europe, 292 UK, 146 Latin America and the Caribbean, and 118 Africa participants.
  • Study metadata: Survey completion dates span 2023-11-22 to 2023-12-22, and extreme completion-time values reflect participants completing the task in multiple sessions.

V.4 Metadata Codebook

The metadata codebook defines identifiers and automated quality-control fields linking participants, conversations, utterances, and text annotations. It records possible misclassifications and incomplete manual checks for several automated flags.

  • Identifiers: User, conversation, and utterance identifiers link participant metadata to the main data and distinguish individual human-message/model-response pairs.The dataset contains 1,500 unique user IDs, 8,011 unique conversation IDs, and 68,371 unique utterance IDs.
  • PII and text-quality fields: The pii_flag is an automated binary flag for personally identifiable information, generated with scrubadub and subject to possible misclassification.The documentation notes that many inspected positives were false positives and that all checked positive human-written texts were reviewed.
  • PII and text-quality fields: Manual review overrode automated PII flags for checked human-written text, but model-generated text was not manually checked.NaN indicates entries that were not manually checked.
  • Language and moderation fields: language_flag records automated language detection using langid, with 59 unique values and possible misclassifications.
  • Language and moderation fields: moderation_flag stores OpenAI moderation API outputs as nested binary flags and probabilities for harm subcategories, which may also be misclassified.
Loading 2404.16019v2…