Chatbot Evaluation on ArenaHard v2
57.4ArenaHard v2 ScoreDeepSeek R1
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DeepSeek R1Input Cost (per M tokens)=0.4, Output Cost (per M tokens)=2, Relative Cost ($)=25x2025.09 | 57.4 | — | — | |
| Qwen3-32B + RLBFF trainingInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 55.6 | — | — | |
| Claude-3.7-Sonnet (Thinking)Input Cost (per M tokens)=3, Output Cost (per M tokens)=15, Relative Cost ($)=188x2025.09 | 54.2 | — | — | |
| o3-miniInput Cost (per M tokens)=1.1, Output Cost (per M tokens)=4.4, Relative Cost ($)=61x2025.09 | 50 | — | — | |
| Qwen3-32B + Baseline BT trainingInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 47.5 | — | — | |
| Qwen3-32BInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 44 | — | — | |
| Base2026.01 | — | 14 | 13.7 | |
| GRPO2026.01 | — | 12 | 10.8 | |
| SDPO2026.01 | — | 12.3 | 11.1 | |
| SFT on self-teacher2026.01 | — | 11.2 | 8.9 |