Chatbot Evaluation on MT-Bench (GPT-4-Turbo Score)
9.5Score (GPT-4-Turbo)Qwen3-32B + RLBFF training
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-32B + RLBFF trainingInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 9.5 | |
| DeepSeek R1Input Cost (per M tokens)=0.4, Output Cost (per M tokens)=2, Relative Cost ($)=25x2025.09 | 9.49 | |
| Qwen3-32B + Baseline BT trainingInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 9.45 | |
| Qwen3-32BInput Cost (per M tokens)=0.018, Output Cost (per M tokens)=0.072, Relative Cost ($)=1x2025.09 | 9.38 | |
| o3-miniInput Cost (per M tokens)=1.1, Output Cost (per M tokens)=4.4, Relative Cost ($)=61x2025.09 | 9.26 | |
| Claude-3.7-Sonnet (Thinking)Input Cost (per M tokens)=3, Output Cost (per M tokens)=15, Relative Cost ($)=188x2025.09 | 8.93 |