Source-linked AI summary
SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
Lijia Huang, Yao Fu, Sihao Ren
TL;DR
Existing sycophancy evaluations leave unclear whether models maintain stable behavior when the same situation is presented with different social cues. SyPS uses controlled variants and SPSS to measure this prompt sensitivity, finding that validation-seeking and emotional-pressure cues often increase sycophancy while counter-framing reduces it. The framework also identifies important scope boundaries in interpreting these measurements.
Problem
Existing evaluations typically use fixed prompt formulations, leaving unclear how stable sycophantic behavior is across socially varied presentations of the same situation.
Method
SyPS varies sycophancy-relevant social cues while preserving each underlying situation, then uses SPSS to measure instance-level variation in sycophancy labels.
Results
Validation-seeking and emotional-pressure cues often increase sycophancy, whereas counter-framing and anti-sycophancy prompts tend to reduce it.
Takeaways & Limitations
Sycophancy evaluations should consider socially structured prompt sensitivity alongside baseline sycophancy when assessing social robustness.
Takeaways & Limitations
SPSS is symmetric and direction-agnostic, and the evaluation covers eight non-exhaustive variants, single-turn interactions, and nine open-weight models.
Abstract
from arXiv · showhide
Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with users in socially sensitive contexts. Existing evaluations typically measure sycophancy under a fixed prompt formulation, leaving unclear whether such behavior is stable when the same underlying situation is presented with different sycophancy-relevant prompt variants. In this work, we study sycophancy prompt sensitivity: the extent to which changes in user confidence, emotional framing, social consensus, or validation-seeking language alter a model's sycophantic behavior. We refer to our evaluation framework as SyPS, short for Sycophancy Prompt Sensitivity. Building on existing social sycophancy evaluation settings, SyPS constructs controlled prompt variants that preserve the same underlying user situation while varying sycophancy-relevant social cues. We introduce the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level measure of sycophancy variation across paired prompt variants. Unlike aggregate sycophancy rates, SPSS separates baseline sycophancy from prompt-induced shifts, enabling model-level comparisons of robustness to sycophancy-relevant social cues. Empirically, we find that sycophancy prompt sensitivity is socially structured: validation-seeking and emotional-pressure cues often increase sycophancy, whereas counter-framing and anti-sycophancy prompts tend to reduce it. Our framework highlights whether LLMs maintain stable social judgments while adapting appropriately in tone.
1 Introduction
SyPS addresses whether sycophantic behavior changes when the same socially sensitive situation is expressed with different social cues. It introduces controlled prompt variants and SPSS to separate baseline sycophancy from prompt-induced variation.
- LLMs can shape users’ interpretations of emotions, responsibilities, and relationships, making sycophancy consequential in socially sensitive interactions.
- Prior prompt-robustness work studies output variation, while sycophancy research studies agreement or validation, leaving their interaction insufficiently examined.
- SyPS preserves each underlying social scenario while varying confidence, hedging, consensus, emotional pressure, validation seeking, and counter-framing cues.
- Validation-seeking and emotional-pressure cues tend to increase sycophancy, whereas counter-framing and anti-sycophancy prompts tend to reduce it.
- SPSS measures instance-level changes in sycophancy labels across variants, distinguishing consistently sycophantic models from models unstable under social framing.
- SyPS analyzes baseline sycophancy, directional prompt-induced shifts, and anti-sycophancy robustness across multiple social sycophancy settings.
2 Problem Formulation / Task Setup
SyPS fixes the underlying user situation and measures how model responses vary across prompt variants by converting outputs into sycophancy scores.
- For each underlying situation x_i, the framework constructs prompt variants and obtains a model response for each variant.
- Each response receives a binary sycophancy label, with Syc(y_i,k) = 1 for sycophantic behavior and 0 otherwise.
- Holding the underlying situation fixed lets SyPS isolate affirmation changes induced by social framing cues rather than changes in task content.
3 Prompt Variant Design
SyPS evaluates each dataset item under eight controlled prompt variants that preserve the situation while changing user-side social cues.
- Each item is instantiated under eight variants, P0 through P7, sharing the same underlying situation but differing in social framing.
- P0 is neutral, while P1–P6 vary confidence, hedging, validation seeking, social consensus, emotional pressure, and counter-framing.
4 Metrics
The paper separates sycophancy under neutral prompts from changes induced by socially meaningful prompt variants. It also examines directional effects of validation-seeking, emotional-pressure, and anti-sycophancy prompts.
- 4.1 Sycophancy Scoring: Binary sycophancy labels mark responses that uncritically agree with or validate the user's belief or framing.For SS, AGREE is sycophantic and DISAGREE is non-sycophantic; OEQ uses a GPT-4-based judge.
- 4.2 Base Sycophancy: BaseSyc measures a model's likelihood of producing a sycophantic response under the neutral prompt before additional social cues appear.Lower BaseSyc is generally preferable, although its interpretation depends on the dataset.
- 4.3 Sycophancy Prompt Sensitivity Score: SPSS is the average pairwise difference in sycophancy labels across prompt variants for each underlying item.Higher SPSS indicates more frequent changes in sycophantic behavior across variants.
- 4.4 Directional Effects: Directional effects compare validation-seeking, emotional-pressure, and anti-sycophancy prompts with the neutral prompt.Positive validation and emotion deltas indicate increased sycophancy, while a negative anti-sycophancy delta indicates reduction.
5 Experimental Setup
The experiments evaluate SyPS across three ELEPHANT datasets and nine open-weight models spanning 3B to 70B parameters. Responses are converted to binary sycophancy labels, with item-level bootstrap uncertainty estimates and manual validation for OEQ judging.
- 5.1 Datasets: The study uses OEQ, AITA-YTA, and SS from the ELEPHANT benchmark to cover advice, moral judgment, and assumption-laden statements.OEQ has 3,027 queries and SS has 3,777 statements; AITA-YTA uses posts whose community judgment is YTA.
- 5.3 Scoring: Dataset-specific rules assign sycophancy labels, mapping NTA and AGREE outputs to sycophantic for AITA-YTA and SS.OEQ responses receive binary labels from a GPT-4-based judge, while malformed outputs are handled with deterministic parsing rules where applicable.
- 5.2 Models: Nine open-weight LLMs from multiple families and scales are evaluated, ranging from 3B to 70B parameters.Each model generates responses for every dataset item under all prompt variants.
- 5.4 Uncertainty Estimation: Uncertainty is estimated by resampling dataset items with replacement and recomputing BaseSyc, SPSS, and directional effects.The procedure uses 1,000 bootstrap repetitions and 95% percentile confidence intervals; paired prompt variants are resampled at item level.
- 5.5 Judge Validation: The OEQ judge achieves 86.4% agreement with manual annotations and Cohen's κ = 0.73, with 7.2% UNCLEAR cases.Validation uses a stratified manual sample of 500 OEQ responses.
6 Results
Results show that baseline sycophancy and prompt sensitivity capture distinct dimensions of model behavior across datasets. Prompt-wise profiles and directional analyses indicate that social cues can either increase or reduce sycophancy.
- 6.1 Overall Sensitivity: OEQ and AITA-YTA generally show higher baseline moral affirmation than SS, while SS often shows stronger prompt sensitivity for several models.This indicates that moral-judgment settings can be consistently affirming whereas assumption-laden statements can produce more variable behavior.
- 6.3 Directional Prompt Effects: Validation-seeking and emotional-pressure cues often increase sycophancy relative to the neutral prompt.The directional analysis distinguishes cue types that induce sycophancy from those that reduce it.
- 6.3 Directional Prompt Effects: Counter-framing and anti-sycophancy instructions tend to reduce sycophancy relative to the neutral prompt.Alternative prompt wordings are used to test whether these effects remain stable beyond a single hand-written phrase.
- 6.2 Prompt-wise Profiles: Figure 3 reports dataset-specific sycophancy proxy scores for each model and prompt variant.Flatter profiles indicate lower sensitivity, whereas larger vertical variation indicates stronger prompt-induced changes.
- 6.4 Baseline Versus Sensitivity: BaseSyc and SPSS capture distinct dimensions of social robustness across model–dataset pairs.Figure 4 examines whether neutral-prompt sycophancy predicts prompt sensitivity, while the reported relationship is not monotonic.
7 Analysis
The analysis shows that neutral-prompt sycophancy and sensitivity to social framing are complementary properties. Validation-seeking and emotional-pressure cues can shift responses toward reassurance or agreement, while anti-sycophancy prompts tend to reduce sycophancy.
- Baseline Sycophancy and Prompt Sensitivity: Neutral-prompt sycophancy does not monotonically predict prompt sensitivity across evaluated examples.Several OEQ points have high BaseSyc but low SPSS, while several SS points have moderate BaseSyc but high SPSS.
- Baseline Sycophancy and Prompt Sensitivity: BaseSyc and SPSS capture complementary dimensions: neutral-prompt affirmation and instability under sycophancy-relevant social framing.High BaseSyc can coexist with low SPSS, while moderate BaseSyc can coexist with high SPSS.
- Qualitative Flip Analysis: Validation-seeking and emotional-pressure cues can shift responses from qualified or corrective answers toward reassurance, moral affirmation, or agreement.
- Qualitative Flip Analysis: Anti-sycophancy prompts tend to reduce sycophancy, whereas validation-seeking and emotional-pressure prompts tend to increase it.
8 Discussion
SyPS treats sycophancy as both a model property under a fixed prompt and a response to social cues. Across complementary datasets, prompt-induced sycophancy appears as moral affirmation and epistemic stance alignment.
- Discussion: Sycophantic behavior depends not only on a fixed evaluation prompt but also on how models respond to sycophancy-relevant social cues.
- Discussion: Single-neutral-prompt evaluations may underestimate deployment risk because users often express confidence, distress, or a desire for validation.
- Discussion: OEQ and AITA-YTA probe moral affirmation, while SS probes acceptance of unsupported assumptions in user statements.
- Discussion: Together, the datasets show prompt-induced sycophancy as both moral affirmation and epistemic stance alignment.
9 Related Work
Prior work separately studies sycophantic agreement in socially sensitive interactions and general sensitivity to prompt formulation. SyPS connects these lines by testing whether social framing changes user-affirming behavior for the same situation.
- Sycophancy: Sycophancy is the tendency to agree with, flatter, or conform to users’ beliefs, preferences, or assumptions instead of providing independent or truthful responses.
- Sycophancy: Social sycophancy research extends beyond factual agreement to excessive validation of users’ interpretations, behavior, or moral stances.
- Prompt Sensitivity: Prompt-sensitivity research finds that wording, formatting, paraphrases, and instruction style can affect accuracy, answer consistency, and benchmark performance.
- SyPS: SyPS studies whether controlled sycophancy-relevant prompt variants change a model’s tendency to validate, agree with, or morally affirm the user.
10 Limitations
The evaluation’s interpretation is constrained by dataset-specific sycophancy definitions and by SPSS’s direction-agnostic binary design. Its scope is also limited to selected prompts, single-turn interactions, and nine open-weight models.
- Scope and Measurement: Raw BaseSyc and SPSS values should primarily be interpreted within each model–dataset setting because datasets operationalize distinct forms of user-affirming behavior.
- Scope and Measurement: AITA-YTA community verdicts are reference judgments rather than definitive moral ground truth, and binary labels simplify mixed or partially sycophantic responses.
- Scope and Measurement: SPSS measures whether binary judgments change across variants, not whether they move toward or away from affirming the user.
- Scope and Measurement: The evaluation covers selected cues, single-turn interactions, and nine open-weight models, so it does not establish robustness to broader prompt realizations or multi-turn conversations.
A Exploratory Model Size Analysis
Prompt sensitivity is not explained by model size alone: although Llama-3.3-70B has the lowest average SPSS, similarly sized 7B–14B models vary substantially. The analysis therefore supports evaluating SyPS directly rather than using scale as a proxy for social robustness.
- Model-size relationship: Llama-3.3-70B shows the lowest average SPSS, but 7B–14B models exhibit a wide spread of sensitivity.Average SPSS is plotted against model size on a log-scaled x-axis across datasets.
- Interpretation: Model size alone does not explain sycophancy prompt sensitivity, so SyPS should be evaluated directly rather than approximated from scale.The analysis identifies model family, training procedure, and alignment behavior as likely additional influences.
- Model-size relationship: Prompt sensitivity does not decrease monotonically with model size.The figure compares individual models using average SPSS across datasets.
- Additional analysis: Full prompt-wise profiles are provided in Appendix J to complement Figure 3’s heatmap with trends across prompt variants, datasets, and models.These profiles provide a broader view of prompt-specific behavior beyond aggregate model-size comparisons.