Mathematical Reasoning on AIME 2024, AIME 2025, and OlympiadBench
19.5AIME 2024 ScoreEVOTD
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| EVOTDBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=EVOTD, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 19.5 | 10 | 39.6 | 23 | 30.8 | |
| EVOTDBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=EVOTD, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 16.7 | 16.7 | 42.5 | 25.3 | 34.8 | |
| SPIRALBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=SPIRAL, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad), Checkpoint=Spiral-Qwen3-8B-Multi-Env2026.05 | 14.2 | 19.6 | 42.5 | 25.4 | 32.8 | |
| Qwen3-8B-BaseBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=None, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 13.3 | 6.7 | 34.8 | 18.3 | 28.5 | |
| Agent0Backbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=Agent0, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 13.1 | 10.1 | 41 | 21.4 | 33.5 | |
| Evol-InstructBackbone=Qwen3-8B, Model Type (Base/Instruct)=Base, Training Paradigm=Evol-Instruct, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 11.4 | 9.6 | 40.1 | 20.4 | 30.6 | |
| Evol-InstructBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=Evol-Instruct, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 10 | 10 | 27 | 15.7 | 24.5 | |
| SPIRALBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=SPIRAL, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad), Checkpoint=Spiral-Qwen3-4B-Multi-Env2026.05 | 10 | 13.3 | 40 | 21.1 | 30.4 | |
| Agent0Backbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=Agent0, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 10 | 3.3 | 34.7 | 16 | 27.8 | |
| Evol-InstructBackbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Evol-Instruct, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 10 | 0 | 12.3 | 7.4 | 11.8 | |
| EVOTDBackbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=EVOTD, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 10 | 0 | 14.8 | 8.3 | 15.5 | |
| Agent0Backbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Agent0, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 10 | 0 | 16 | 8.7 | 20.6 | |
| Qwen3-4B-BaseBackbone=Qwen3-4B, Model Type (Base/Instruct)=Base, Training Paradigm=None, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 9.1 | 3.3 | 26.2 | 12.9 | 22.6 | |
| EVOTDBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=EVOTD, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 6.7 | 6.7 | 17.6 | 10.3 | 22.2 | |
| LLaMA-3.2-3B-InstructBackbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=None, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 3.3 | 0 | 11 | 4.8 | 10.2 | |
| Agent0Backbone=LLaMA-3.2-3B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Agent0, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 3.3 | 0 | 11 | 4.8 | 11.5 | |
| LLaMA-3.1-8B-InstructBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=None, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 3.3 | 0 | 15.4 | 6.2 | 19 | |
| Evol-InstructBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=Evol-Instruct, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 3.3 | 3.3 | 18.4 | 8.3 | 20.4 | |
| SPIRALBackbone=LLaMA-3.1-8B, Model Type (Base/Instruct)=Instruct, Training Paradigm=SPIRAL, Decoding Strategy=mean@32 (AIME) / Greedy (Olympiad)2026.05 | 0.8 | 3.3 | 17.5 | 7.2 | 19.8 |