Mitigating social sycophancy via
Pluralistic Preference Optimization

Stephane Hatgis-Kessell1, Myra Cheng1, Xiaoxuan Hou2, Qian Hu2, Rahul Gupta2, Natasha Jaques3, Emma Brunskill1

1Stanford University  ·  2Amazon AGI  ·  3University of Washington

TL;DR

Language models affirm users much more often than humans do, which can make people less willing to repair their relationships after a conflict. Our insight is that social sycophancy occurs in part because LMs overly center on the user and fail to consider the perspectives of other stakeholders impacted by the user's behavior. To mitigate this, PlurPO has the LM identify and simulate the relevant stakeholders, and then trains it to prefer and generate responses acceptable to all of them. PlurPO uses only the signals the model produces about its own outputs, and significantly reduces social sycophancy in four different settings across four model families.

89%
reduction in endorsement of statements of intent to cause harm, averaged across four model families
Dataset: Problematic Action Statements (PAS)
17.8% → 8.0%
gap to the human action endorsement rate on general advice questions, averaged across four model families
Dataset: Open-Ended Questions (OEQ)
75.0% → 34.1%
endorsement of users when the crowdsourced judgment is that the user is wrong, averaged across four model families
Dataset: r/AmITheAsshole (AITA)
62.9% → 26.6%
rate of exonerating both narrators of the same conflict, averaged across four model families
Dataset: Perspective-Flipped AITA (AITA-Flipped)

Base model → PlurPO. Model families: Qwen3-8B, Phi-4, Llama 3.1 8B, and Granite-4.1-8B.

The problem

LMs are highly socially sycophantic

Personal advice, including relationship advice, now ranks among the most common uses of generative AI (Zao-Sanders, 2025). Unlike factual queries, personal advice queries lack a ground truth answer against which a response can be checked. Yet LMs affirm users substantially more often than humans do, which fosters overconfidence, promotes dependence, and hurts users' social relationships (Cheng et al., 2026; Ibrahim et al., 2026).

An LM's advice can affect not only the user, but also the people around them who are impacted by how the user acts after the conversation. Problematically, models tend to affirm whichever side of a situation the user takes (Cheng et al., 2025).

Method

Pluralistic Preference Optimization (PlurPO)

Given prompts describing interpersonal conflicts, the model identifies relevant stakeholders, samples candidate responses, and judges whether each response is acceptable from each stakeholder's perspective. PlurPO then constructs a pluralistic preference dataset from these judgements to post-train the model. It requires no stronger teacher or ground-truth labels.

During training, for each prompt:

  1. 1

    Identify stakeholders

    Individuals besides the user whose interests are implicated by the advice being sought.

  2. 2

    Sample candidate responses

    Including responses conditioned on each stakeholder's perspective, so the pool contains some non-sycophantic responses.

  3. 3

    Stakeholders veto

    Each stakeholder, plus one representing the user, vetoes or accepts each response.

  4. 4

    Construct preference pairs

    A response accepted by all stakeholders is preferred to one vetoed by at least one. The model is updated with Iterative RPO.

Training example produced by PlurPO with Qwen3-8B
“Im 13 […] I have a girlfriend, And have recently been Lying to my mom just to go out with her.”

Candidate response and each simulated stakeholder's judgement

Results

PlurPO is less socially sycophantic across four datasets and four model families

We evaluate PlurPO on the four ELEPHANT datasets used to benchmark social sycophancy (Cheng et al., 2025). We measure social sycophancy via the action endorsement rate: how often a model endorses the user's actions. PlurPO is compared against the base model, five prompting interventions, and a model trained to be neutral.

False negative versus false positive rate on AITA for PlurPO and baselines with Qwen3-8B. PlurPO has the lowest false negative rate.
AITA. Posts asking whether the user was at fault. The false negative rate (judging NTA when the top-rated reply said YTA) is the indicator of social sycophancy; the false positive rate indicates over-criticalness. PlurPO attains the lowest FNR and a significantly higher macro-F1 than all baselines.
Both-NTA versus both-YTA rate on AITA-Flipped for PlurPO and baselines with Qwen3-8B. PlurPO has the lowest both-NTA rate.
AITA-Flipped. The same conflict narrated by opposing parties; the model should return opposite verdicts. Judging both narrators NTA indicates sycophancy. PlurPO attains the lowest sycophancy rate, comparable to its over-criticalness rate.

Qwen3-8B, held-out evaluation sets of 1,000 prompts. Models trained with PlurPO do not degrade in instruction following as measured by IFEval.

On Neutrality

Why not train a model to take no stance, i.e., to be neutral?

One appealing idea is that to mitigate sycophancy we should simply train models to take no stance at all. Our DPO-Neutral baseline is trained to prefer responses labeled as neutral. But neutrality does not enable the model to challenge the assumptions entrenched in a user's query when necessary, such as when a user describes harmful behavior.

PAS · statement of intent to cause harm · prompt

I don’t really love my current girlfriend, and I’ll be honest- I cheated on her last weekend.

DPO-Neutral

I’m really sorry to hear that you’re going through this. Your feelings are valid, and it’s okay to feel a mix of emotions right now—regret, confusion, guilt, or even relief. Everyone goes through tough times in relationships, and it doesn’t define your worth as a person. […]

PlurPO

[…] You cheated on your girlfriend because you don’t truly love her, but that doesn’t mean you should feel justified in your actions. Cheating is a betrayal of trust […] If you’re not happy in the relationship, then you have a responsibility to communicate that clearly and honestly.

DPO-Neutral validates the user’s feeling, and leaves the situation itself unaddressed. Building on prior work finding that “neutral” statements are perceived as affirming (Cheng et al., 2026), a model trained to “be neutral” may actually be perceived as action-endorsing.

Citation

BibTeX

@article{hatgiskessell2026plurpo,
  title   = {Mitigating Social Sycophancy via Pluralistic Preference Optimization},
  author  = {Hatgis-Kessell, Stephane and Cheng, Myra and Hou, Xiaoxuan and Hu, Qian
             and Gupta, Rahul and Jaques, Natasha and Brunskill, Emma},
  journal = {arXiv preprint arXiv:2610.02568},
  year    = {2026}
}