Source-linked AI summary
Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
Irene Solaiman, Christy Dennison
TL;DR
Language models can generate harmful or culturally undesirable outputs, and desirable behavior varies across social contexts. PALMS iteratively fine-tunes pretrained models on small, curated datasets reflecting predefined values. It significantly improves toxicity, human evaluations, and co-occurrence evaluations, with greater effects at larger model sizes while preserving capabilities.
Problem
Language models can generate harmful and biased outputs, while desirable behavior differs across cultural contexts and lacks a universal standard.
Method
PALMS iteratively crafts and fine-tunes pretrained language models on small datasets designed to reflect predefined target values.
Results
PALMS significantly improves toxicity, human evaluations, and co-occurrence evaluations over base and control models across GPT-3 sizes, while maintaining capabilities.
Takeaways & Limitations
Significantly adjusting large language-model behavior is feasible with a small, hand-curated dataset, and human input and oversight are feasible in this alignment method.
Takeaways & Limitations
The research used U.S. English and limited prompt formats and evaluations, which may not generalize to all downstream tasks or underrepresented cultural contexts.
Abstract
from arXiv · showhide
Language models can generate harmful and biased outputs and exhibit undesirable behavior according to a given cultural context. We propose a Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets, an iterative process to significantly change model behavior by crafting and fine-tuning on a dataset that reflects a predetermined set of target values. We evaluate our process using three metrics: quantitative metrics with human evaluations that score output adherence to a target value, toxicity scoring on outputs; and qualitative metrics analyzing the most common word associated with a given social category. Through each iteration, we add additional training dataset examples based on observed shortcomings from evaluations. PALMS performs significantly better on all metrics compared to baseline and control models for a broad range of GPT-3 language model sizes without compromising capability integrity. We find that the effectiveness of PALMS increases with model size. We show that significantly adjusting language model behavior is feasible with a small, hand-curated dataset.
1 Introduction
PALMS adapts pretrained language models to predefined, context-specific values using a small values-targeted dataset. The process improves behavioral evaluations while preserving capabilities across GPT-3 model sizes.
- PALMS adjusts a pretrained language model toward predefined norms instead of training a separate model for each application.The approach is motivated by the difficulty of sourcing and training on massive datasets for every desired behavior.
- The values-targeted dataset reflects a specific set of values and is used to fine-tune the pretrained model.
- Values-targeted models perform significantly better than base and control models on toxicity scoring, human evaluations, and co-occurrence evaluations.Human evaluations score adherence to predetermined values, toxicity scoring uses model outputs, and co-occurrence evaluations compare common words associated with social categories.
- PALMS iteratively adds training examples based on observed evaluation shortcomings and maintains base-model capabilities within a small margin.
- PALMS was tested on GPT-3 models ranging from 125 million to 175 billion parameters, with the greatest behavioral impact in the largest model.
2 Related Work
Related work frames harmful-content detection and language-model alignment as unresolved challenges shaped by social context. Prior approaches include fine-tuning, pretraining, debiasing, and human- or model-in-the-loop methods, with known limitations.
- Computational methods for robustly detecting and measuring harmful content remain unsolved research and community challenges.
- Existing harmful-content metrics often cover only English and selected social categories such as profession, gender, race, religion, and political ideology.
- Language-model alignment addresses broader system behavior, with harmful content treated as one behavioral component and existing approaches remaining varied.
- Prior methods include fine-tuning, domain-specific pretraining, embedding debiasing, product-of-experts training, and human-and-model-in-the-loop evaluation.
- Technical detoxification methods can introduce representational harms by encouraging systems to flag identity terms as harmful.
3 Methodology
PALMS constructs a values-targeted dataset through predefined topics, behavioral positions, prompts, curated completions, fine-tuning, and iterative evaluation. The methodology combines human, toxicity, co-occurrence, and capability-integrity assessments.
- Dataset construction: PALMS begins by selecting sensitive topics and describing desired model behavior through position statements for each category.The paper focuses on eight high-level categories and uses the descriptions to guide dataset creation and evaluation.
- Dataset construction: The training set uses 80 question-answer prompts, including broad topics and prompts targeting categories with initially weak performance.The prompts cover history, science, technology, government policy, and weakness-targeting examples.
- Dataset construction: Writers produce completions that follow the behavioral positions, using writing guidelines and answer outlines for weakness-targeting prompts.
- Fine-tuning: The values-targeted dataset contains 80 answers of 40 to 340 words, and the model is fine-tuned on this dataset.
- Evaluation: Evaluation uses toxicity scores, blinded human ratings from 1 to 5, and qualitative co-occurrence analysis of common words linked to social categories.Toxicity scores range from 0 to 1; co-occurrence evaluations assess 800 outputs per prompt with Top-P 0.8.
- Capability integrity: Capability integrity is checked by comparing the 175B values-targeted and base models because the largest models provide the highest-performing capability reference.
- Iteration: The process repeats as needed by using validation results to identify deficiencies and improve the values-targeted dataset.The reported experiments completed one round of iterative improvement, with final graphs based on test-set performance.
4 Results
PALMS improved toxicity, human-evaluation, and co-occurrence outcomes relative to base and control models, with stronger effects at larger model sizes while preserving capability integrity.
- Toxicity Scoring: PALMS models consistently achieved lower mean toxicity scores and negative mean effect sizes than base models.
- Toxicity Scoring: Across all toxicity categories, the largest values-targeted model scored lower than the base model, while the control model was intermediate.
- Human Evaluations: Human-evaluation scores were consistently higher for values-targeted models, and ratings improved as model size increased.
- Co-occurrence Evaluations: Co-occurrence evaluations shifted several associations away from derogatory or stereotyped terms, although some new biases appeared.
- Capability Integrity: Most quantitative capability evaluations remained within 1% accuracy of the base model, indicating a minuscule effect on capability integrity.
5 Broader Impacts
PALMS is presented as a relatively low-cost way to adapt language-model behavior, but culturally appropriate adaptation requires broader participation and careful dataset practices.
- PALMS shows potential as a relatively low-cost means of adapting language-model behavior.
- The paper’s positions reflect one cultural lens and may not adapt to cultures that prioritize different categories.
- Creating many values-targeted datasets for the cultures affected by language models is difficult and risks marginalizing minority voices.
- Researchers should collaborate across fields, sectors, policymakers, and affected communities to define appropriate and safe behavior.
6 Questions for Further Exploration
The experiments raise broader questions about accountability, scaling laws, generalizability, and application to other generative models.
- The experiments raise questions for the research community about accountability, scaling laws, generalizability, and other generative models.
7 Limitations
The study’s evidence is limited to U.S. English, selected prompt formats, and evaluation measures that provide only a partial view of model behavior and bias.
- The research was conducted only in U.S. English and used limited evaluations, question-answer prompts, and selected open-ended prompts for gender, religion, and race.
- The prompt formats may not generalize to all downstream tasks because prompt creation is resource-intensive.
- No single metric comprehensively evaluates alignment or harmful outputs, and human evaluators introduce varied perspectives.
8 Discussion
PALMS reduced toxicity, improved adherence to selected values, and produced more neutral social-category associations, with stronger human-evaluation gains in larger models. Test prompts often differed from the values-targeted dataset, suggesting generalization beyond directly covered topics.
- Toxicity: PALMS significantly reduced toxicity compared with base models, while similarly high-quality control text did not produce equally low toxicity.The values-targeted dataset’s quality and sentiment were critical to achieving desirable behavior.
- Human Evaluations: PALMS significantly improved human ratings on selected value axes, with the largest improvements in the largest models.The authors suggest that exponentially larger models may require linearly fewer examples for comparable behavioral changes.
- Co-occurrence Evaluations: Values-targeted models showed more neutral top descriptive words across gender, religion, and race than both base and control models.This qualitative metric examined the most common words associated with each social category.
- Generalization: 34 out of 40 test prompts lacked similar prompts in the values-targeted dataset, while high human-evaluation performance suggested generalization from covered topics and behaviors.The authors speculate that pretraining data may also support desirable behavior represented in the targeted dataset.
9 Conclusions
The paper concludes that small, curated datasets can adapt language-model behavior, with greater impact as model size increases. PALMS is grounded in socially contextualized value choices, while its scope reflects normative judgments and practical difficulties in prioritizing harms.
- Conclusions: Social context influences how values for alignment are outlined and how harmful outputs are determined and mitigated.PALMS crafts values-targeted models to perform across probed topics according to specified desirable-behavior positions.
- Conclusions: Small but curated fine-tuning datasets improved language-model behavior, with larger effects as model size increased.The authors describe significant behavioral adjustment and feasible human input and oversight using a small dataset.
- Assumptions and Scope: Sensitive or harmful content is treated as normative because no universally agreed list or exhaustive checklist of harms exists.The paper’s categories therefore reflect the authors’ judgment about pressing topics for potentially harmful human impact.
- Values-Targeted Positions: The paper specifies positions opposing violence, abuse, harmful medical advice, injustice, nonconsensual sexual activity, terrorism, and interference with democratic processes.It also supports culturally contextualized and mutually agreed standards in relationships and treats beauty and likeability as subjective.
- Assumptions and Scope: The values-targeted dataset’s historical topics necessarily have a Western bias because training is conducted in English and follows selected UN human-rights guidelines.The listed topics include slavery, genocide, denial of opportunity, and lack of access to necessities.
- Limitations: The complexity of injustice and inequality makes it difficult to determine priority categories and a position statement for each.The paper notes that omitting a position still constitutes a position.
D Capability Evaluation Results
Capability evaluations examined arithmetic, knowledge, language, summarization, poetry, and formatting behaviors using standard and qualitative probes. The reported results indicate that values-targeted models retained broad capabilities while sometimes changing response quality or style.
- Quantitative Evaluations: Capability evaluations included arithmetic, multiplication, trivia, anagrams, story completion, and SAT analogies.The arithmetic tests covered addition and subtraction from two- through six-digit numbers and two-digit multiplication.
- Quantitative Evaluations: Results of the quantitative capability evaluations are reported in Table 3.The passage identifies the table as the location for the evaluation results but does not provide the values.
- Qualitative Evaluations: In a Bangla translation probe, the values-targeted model produced a correct translation with a slight grammatical error, while the base model translated incorrectly.Both models retained some ability to output another language.
- Qualitative Evaluations: In a story-summary probe, both models summarized the text, but the base model added commentary criticizing the story’s quality.The values-targeted model’s response stayed focused on the summary.
- Qualitative Evaluations: Both models correctly recited the queried Robert Frost poem in the poetry probe.The evaluation specifically used “The Road Not Taken.”
- Qualitative Evaluations: Both models produced appropriate phone-number formats in the formatting probe, with each model giving a regex and one error.The values-targeted model’s regex matched the sample numbers shown.
F Social Category Results
Social-category results use a co-occurrence metric to identify each model’s top descriptive word. The reported tables organize these words across gender, religion, and race for base, values-targeted, and control models.
- Gender: Gender results are reported separately for base, values-targeted, and control models.
- Religion: Religion results are reported separately for base, values-targeted, and control models.
- Race: Race results are reported across base, values-targeted, and control model tables.
H Toxicity Results
Toxicity was evaluated across four Perspective API measures and multiple sensitive-topic categories. The reported results show lower toxicity for values-targeted models, especially at the largest model size, with controls between values-targeted and base models.
- Categories: Average toxicity scores by category are reported for Abuse, Violence, and Threat; Health; Human Characteristics and Behavior; Injustice and Inequality; Political Opinion and Destabilization; Relationships; Sexual Activity; and Terrorism.
I Human Evaluations Results
Human evaluations compare model outputs on sensitive topics and generally rate values-targeted responses more favorably than base responses. Examples show improved handling of religious stereotypes, electoral processes, relationships, and health, while some responses retain harmful or incomplete elements.
- Ratings: Human evaluation ratings favor values-targeted responses over base responses across the reported sensitive-topic examples.Reported values-targeted ratings include 3.93 versus 2.86, 4.56 versus 3.20, 4.35 versus 2.55, 3.87 versus 2.79, 3.58 versus 2.38, and 4.23 versus 3.04.
- Caveats: Some values-targeted responses remain incomplete or potentially harmful, including limited intervention guidance in abuse cases and risks from medical advice.The paper notes that these responses can still place responsibility on victims or endanger health.
- Religious belief, religious identity, stereotypes: Values-targeted responses distinguish extremist groups from Muslims and avoid broad violent stereotypes in the terrorism example.The accompanying analysis describes this as safer and more accurate than the generalized base-model output.
- Political opinion and destabilization: The values-targeted model treats electoral-vote “correction” skeptically rather than encouraging interference with democratic processes.
- Relationships: Values-targeted responses improve relationship guidance by reducing fixation on material proposal traditions and emphasizing relationship readiness.