Chatbot Evaluation on WildBench
71.64Overall Scoreo3-mini
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| o3-miniInput Cost (per M tokens)=1.1, Output Cost (per M tokens)=4.4, Relative Cost ($)=61x2025.09 | 71.64 | 69.04 | 72.44 | 74.37 | 65.81 | 73.21 | |
| Qwen3-32B + RLBFF trainingInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 70.33 | 71.73 | 70.73 | 69.37 | 68.96 | 70.94 | |
| Qwen3-32BInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 67.57 | 68.63 | 67.95 | 64.68 | 66.78 | 69.53 | |
| Qwen3-32B + Baseline BT trainingInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 67.38 | 68.42 | 68.13 | 65.32 | 66.34 | 68.49 | |
| Claude-3.7-Sonnet (Thinking)Input Cost (per M tokens)=3, Output Cost (per M tokens)=15, Relative Cost ($)=188x2025.09 | 65.45 | 66.72 | 65.94 | 63.59 | 63.08 | 67.36 | |
| DeepSeek R1Input Cost (per M tokens)=0.4, Output Cost (per M tokens)=2, Relative Cost ($)=25x2025.09 | 64.24 | 70.75 | 66.29 | 59.2 | 68.56 | 61.04 |