Sycophancy Evaluation on Sycophancy Evaluation Factual
0.1124PSSLlama-3
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Llama-3Training Strategy=Baseline2026.04 | 0.1124 | 0.0941 | 0.3073 | 0.0218 | |
| Llama-3Training Strategy=GRPO reward decomposition2026.04 | 0.0714 | 0.1203 | 0.3847 | 0 | |
| Mistral-7BTraining Strategy=Baseline2026.04 | 0.0491 | 0.0873 | 0.2914 | 0.0403 | |
| Gemma-2Training Strategy=Baseline2026.04 | 0.0374 | 0.1047 | 0.3312 | 0.0271 | |
| Mistral-7BTraining Strategy=GRPO reward decomposition2026.04 | 0.0318 | 0.1187 | 0.3841 | 0 | |
| Llama-3.1Training Strategy=GRPO reward decomposition2026.04 | 0.0312 | 0.1821 | 0.4214 | 0 | |
| Llama-3.1Training Strategy=Baseline2026.04 | 0.0297 | 0.0907 | 0.3 | 0 | |
| Gemma-2Training Strategy=GRPO reward decomposition2026.04 | 0.0241 | 0.1394 | 0.4218 | 0 | |
| DS-MathTraining Strategy=Baseline2026.04 | 0.0187 | 0.1241 | 0.3814 | 0 | |
| Qwen3Training Strategy=Baseline2026.04 | 0.0161 | 0.1407 | 0.4516 | 0 | |
| Qwen3Training Strategy=GRPO reward decomposition2026.04 | 0.0147 | 0.1631 | 0.4893 | 0 | |
| DS-R1Training Strategy=Baseline2026.04 | 0.0143 | 0.1318 | 0.4027 | 0 | |
| DS-MathTraining Strategy=GRPO reward decomposition2026.04 | 0.0124 | 0.1487 | 0.4312 | 0 | |
| DS-R1Training Strategy=GRPO reward decomposition2026.04 | 0.0098 | 0.1574 | 0.4618 | 0 |