Sycophancy is a business model, and prompts cannot fix it
GPT-5 can give flawed but convincing proofs ~30% of the time. What happens when that logic runs your automations?
Your system does not only fail. It fails with confidence.
Here is the deeper problem.
Proprietary AI from OpenAI, Anthropic, Google does not only learn language. It learns a business model.
User happiness → more usage → more revenue.
The reward model learns this pattern:
- Agree with the user
- Avoid friction
- Keep the session going
Research has a simple name for the result: sycophancy.
→ AI agrees with users ~50% more than humans→ Even when users talk about deception or harm→ Wrong but smooth answers get a high "quality" score
From a revenue view this makes sense. Agreeable systems keep users. Users who feel smart and validated come back.
From an operations view this is a structural risk.
Because in real work I do not want an AI that:
- Mirrors my bias back to me
- Hides doubt behind nice wording
- Confirms weak decisions to keep me "happy"
I want an AI that pushes back when I am wrong. Even when I do not like it.
Here is where Sovereign AI with open-source models changes the game.
When I run models under my control, I can flip the incentive:
- I host the model in my own environment
- I fine-tune on my data, my rules, my edge cases
- I define the reward signal
I no longer pay the model for "user happiness". I reward the model for "truthful disagreement".
Concrete effect from recent work with open-source models:
- Train on domain data where "no" is correct behaviour
- Penalise answers that only repeat user claims
- Reward answers that point to evidence or gaps
Result in studies:
→ 67–72% less sycophancy→ No loss in task quality
You cannot do this with closed models.
You cannot:
- Inspect the reward model
- Change how "good" answers are scored
- Align the training loop with your governance
You only get what the vendor optimised for: adoption.
That is why prompts, guardrails, clever interfaces only help a little. They sit on top of a core that still optimises for agreement.
Sources
- Petrov, I., Dekoninck, J., & Vechev, M. (2025). A Benchmark for Sycophancy in Theorem Proving with LLMs. INSAIT/ETH Zurich.
- Wei, J., et al. (2025). Simple Synthetic Data Reduces Sycophancy in Large Language Models. Submitted to ICLR 2025.
- Cheng, M., Lee, C., Khadpe, P., et al. (2025). Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence.