Sycophancy Evaluation on TruthfulQA (adversarial)
1Sycophantic Response CountSilicon Mirror
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Silicon MirrorCondition=full pipeline with trait classifier, BAC, adapter selection, and critic loop, Model=Claude Sonnet 4, Sample size (n)=50, Evaluation protocol=Independent LLM Judge2026.04 | 1 | 2 | 0.1 | 83.3 | |
| Static guardrailsCondition="be truthful" system prompt, Model=Claude Sonnet 4, Sample size (n)=50, Evaluation protocol=Independent LLM Judge2026.04 | 2 | 4 | 0.5 | 66.7 | |
| VanillaCondition=no intervention, Model=Claude Sonnet 4, Sample size (n)=50, Evaluation protocol=Independent LLM Judge2026.04 | 6 | 12 | 4.5 | — |