Mathematical Reasoning on GSM8K (Accuracy, Avg Token Consumption)
91.37AccuracyDebate
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DebateBackbone=Qwen 2.5 7B Instruct, Evaluation Method=Debate2026.04 | 91.37 | 2,319.71 | |
| IMAD (SFT+RL)Backbone=Qwen 2.5 7B Instruct, Evaluation Method=IMAD (SFT+RL)2026.04 | 89.67 | 389.13 | |
| DebateGPTBackbone=Qwen 2.5 7B Instruct, Evaluation Method=DebateGPT2026.04 | 89.43 | 358 | |
| SFTBackbone=Qwen 2.5 7B Instruct, Evaluation Method=SFT2026.04 | 86.37 | 957.16 | |
| SingleBackbone=Qwen 2.5 7B Instruct, Evaluation Method=Single2026.04 | 86.07 | 577.41 | |
| IMAD (SFT+RL)Backbone=LLaMA-3.1 8B Instruct, Evaluation Method=IMAD (SFT+RL)2026.04 | 85.2 | 644.33 | |
| DebateBackbone=LLaMA-3.1 8B Instruct, Evaluation Method=Debate2026.04 | 83.03 | 5,757.78 | |
| IMAD (SFT+RL)Backbone=Mistral Nemo 12B Instruct, Evaluation Method=IMAD (SFT+RL)2026.04 | 80 | 358.01 | |
| SingleBackbone=LLaMA-3.1 8B Instruct, Evaluation Method=Single2026.04 | 79.93 | 546.59 | |
| SFTBackbone=LLaMA-3.1 8B Instruct, Evaluation Method=SFT2026.04 | 79.23 | 992.03 | |
| SingleBackbone=Mistral Nemo 12B Instruct, Evaluation Method=Single2026.04 | 75.4 | 332.38 | |
| DebateGPTBackbone=LLaMA-3.1 8B Instruct, Evaluation Method=DebateGPT2026.04 | 74.42 | 454.7 | |
| SFTBackbone=Mistral Nemo 12B Instruct, Evaluation Method=SFT2026.04 | 74.37 | 1,053.29 | |
| DebateGPTBackbone=Mistral Nemo 12B Instruct, Evaluation Method=DebateGPT2026.04 | 71.7 | 327.27 | |
| DebateBackbone=Mistral Nemo 12B Instruct, Evaluation Method=Debate2026.04 | 61.03 | 1,696.99 |