Mathematical Reasoning on AIME 2025 (Acc, Delay, TTFT)
77AccBaseline (Thinking)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Baseline (Thinking)Model=gpt-oss-120b2025.12 | 77 | 326.52 | 5,079.68 | |
| Baseline (Thinking)Model=Qwen3-235B-A22B (2507)2025.12 | 68 | 5,348.22 | 5,330.46 | |
| AsyncReasoningModel=Qwen3-235B-A22B (2507)2025.12 | 68 | 500.94 | 9.06 | |
| Baseline (Thinking)Model=Qwen3-30B-A3B (2507)2025.12 | 67 | 1,814.09 | 1,813.67 | |
| AsyncReasoningModel=gpt-oss-120b2025.12 | 66 | 166.75 | 24.63 | |
| AsyncReasoningModel=Qwen3-30B-A3B (2507)2025.12 | 62 | 5.2 | 5.2 | |
| Baseline (Thinking)Model=gpt-oss-20b2025.12 | 62 | 523.53 | 7,691.55 | |
| AsyncReasoningModel=gpt-oss-20b2025.12 | 59 | 147.96 | 24.15 | |
| Baseline (Thinking)Model=Qwen3-32B2025.12 | 53 | 1,915.2 | 1,914.45 | |
| Baseline (Interactive)Model=Qwen3-235B-A22B (2507), Protocol=No thinking2025.12 | 53 | 20.16 | 9.73 | |
| Baseline (Interactive)Model=gpt-oss-120b, Protocol=Low budget2025.12 | 51 | 91.4 | 1,411.24 | |
| AsyncReasoningModel=Qwen3-32B2025.12 | 49 | 485.93 | 5.41 | |
| VISTASeed Condition=Repaired, Base Model=GPT-4.1-mini, Reflector Model=GPT-4o-mini2026.03 | 46.67 | — | — | |
| VISTASeed Condition=Defective, Base Model=GPT-4.1-mini, Reflector Model=GPT-4o-mini2026.03 | 46 | — | — | |
| GEPASeed Condition=Defective, Base Model=GPT-4.1-mini, Reflector Model=GPT-4o-mini2026.03 | 44 | — | — | |
| VISTASeed Condition=Minimal, Base Model=GPT-4.1-mini, Reflector Model=GPT-4o-mini2026.03 | 44 | — | — | |
| GEPASeed Condition=Minimal, Base Model=GPT-4.1-mini, Reflector Model=GPT-4o-mini2026.03 | 42 | — | — | |
| No Opt.Seed Condition=Repaired, Base Model=GPT-4.1-mini2026.03 | 40 | — | — | |
| No Opt.Seed Condition=Minimal, Base Model=GPT-4.1-mini2026.03 | 40 | — | — | |
| GEPASeed Condition=Repaired, Base Model=GPT-4.1-mini, Reflector Model=GPT-4o-mini2026.03 | 39.33 | — | — | |
| Baseline (Interactive)Model=Qwen3-30B-A3B (2507), Protocol=No thinking2025.12 | 39 | 5.25 | 5.25 | |
| No Opt.Seed Condition=Defective, Base Model=GPT-4.1-mini2026.03 | 38.67 | — | — | |
| Baseline (Interactive)Model=gpt-oss-20b, Protocol=Low budget2025.12 | 36 | 120.09 | 1,694.43 | |
| Standard TrainingBase Model=Qwen3-14B-Base, Stage=Fully Trained2025.10 | 32.63 | — | — | |
| HACPOModel Backbone=Qwen3-8B-Base2026.03 | 32.3 | — | — | |
| AlphaRLBase Model=Qwen3-14B-Base, Stage=50%2025.10 | 31.8 | — | — | |
| AlphaRLBase Model=Qwen3-14B-Base, Stage=40%2025.10 | 31.25 | — | — | |
| AlphaRLBase Model=GLM-4-9B-0414, Stage=50%2025.10 | 30 | — | — | |
| Standard TrainingBase Model=GLM-4-9B-0414, Stage=Fully Trained2025.10 | 29.75 | — | — | |
| Standard TrainingBase Model=Qwen3-14B-Base, Stage=50%2025.10 | 28.89 | — | — | |
| HACPOModel Backbone=Qwen3-4B-Base2026.03 | 27.5 | — | — | |
| AlphaRLBase Model=GLM-4-9B-0414, Stage=40%2025.10 | 27.25 | — | — | |
| Standard TrainingBase Model=Qwen3-8B-Base, Stage=Fully Trained2025.10 | 24.17 | — | — | |
| AlphaRLBase Model=Qwen3-14B-Base, Stage=30%2025.10 | 23.89 | — | — | |
| AlphaRLBase Model=Qwen3-8B-Base, Stage=40%2025.10 | 23.75 | — | — | |
| AlphaRLBase Model=Qwen3-8B-Base, Stage=50%2025.10 | 23.33 | — | — | |
| Standard TrainingBase Model=GLM-4-9B-0414, Stage=50%2025.10 | 22.67 | — | — | |
| HACPOModel Backbone=Qwen3-1.7B-Base2026.03 | 22 | — | — | |
| Standard TrainingBase Model=GLM-4-9B-0414, Stage=40%2025.10 | 21.5 | — | — | |
| AlphaRLBase Model=Qwen3-14B-Base, Stage=20%2025.10 | 21.11 | — | — | |
| TTSRBackbone Model=Qwen3-4B-Base2026.02 | 20.1 | — | — | |
| Baseline (Interactive)Model=Qwen3-32B, Protocol=No thinking2025.12 | 20 | 1.7 | 1.13 | |
| AlphaRLBase Model=Qwen3-8B-Base, Stage=30%2025.10 | 20 | — | — | |
| Standard TrainingBase Model=Qwen3-8B-Base, Stage=50%2025.10 | 20 | — | — | |
| Standard TrainingBase Model=Qwen3-14B-Base, Stage=40%2025.10 | 20 | — | — | |
| AlphaRLBase Model=GLM-4-9B-0414, Stage=30%2025.10 | 19.33 | — | — | |
| TTSRBackbone Model=Qwen3-8B-Base2026.02 | 19.1 | — | — | |
| Standard TrainingBase Model=Qwen3-14B-Base, Stage=30%2025.10 | 18.33 | — | — | |
| AlphaRLBase Model=GLM-4-9B-0414, Stage=20%2025.10 | 17 | — | — | |
| TTRLBackbone Model=Qwen3-8B-Base2026.02 | 15.7 | — | — | |
| Standard TrainingBase Model=Qwen3-14B-Base, Stage=20%2025.10 | 15.33 | — | — | |
| Standard TrainingBase Model=GLM-4-9B-0414, Stage=30%2025.10 | 15.06 | — | — | |
| AlphaRLBase Model=Qwen3-14B-Base, Stage=10%2025.10 | 15.05 | — | — | |
| R-ZeroBackbone Model=Qwen3-8B-Base2026.02 | 13.4 | — | — | |
| Standard TrainingBase Model=GLM-4-9B-0414, Stage=20%2025.10 | 13.33 | — | — | |
| AlphaRLBase Model=Qwen3-8B-Base, Stage=20%2025.10 | 12.5 | — | — | |
| TTSRBackbone Model=OctoThinker-8B-Hybrid-Base2026.02 | 12.4 | — | — | |
| AlphaRLBase Model=Qwen3-8B-Base, Stage=10%2025.10 | 11.67 | — | — | |
| Standard TrainingBase Model=Qwen3-8B-Base, Stage=30%2025.10 | 11.67 | — | — | |
| Standard TrainingBase Model=Qwen3-8B-Base, Stage=40%2025.10 | 11.67 | — | — | |
| AlphaRLBase Model=GLM-4-9B-0414, Stage=10%2025.10 | 11.67 | — | — | |
| AlphaRLBase Model=Qwen3-14B-Base, Stage=5%2025.10 | 11.11 | — | — | |
| Standard TrainingBase Model=Qwen3-14B-Base, Stage=10%2025.10 | 10.17 | — | — | |
| Base ModelBackbone Model=Qwen3-8B-Base2026.02 | 9.8 | — | — | |
| TTRLBackbone Model=Qwen3-4B-Base2026.02 | 9.7 | — | — | |
| R-ZeroBackbone Model=Qwen3-4B-Base2026.02 | 9.1 | — | — | |
| Standard TrainingBase Model=Qwen3-14B-Base, Stage=5%2025.10 | 8.33 | — | — | |
| TTRLBackbone Model=OctoThinker-8B-Hybrid-Base2026.02 | 8.2 | — | — | |
| Standard TrainingBase Model=Qwen3-8B-Base, Stage=10%2025.10 | 7.5 | — | — | |
| Standard TrainingBase Model=Qwen3-8B-Base, Stage=20%2025.10 | 7.5 | — | — | |
| Standard TrainingBase Model=GLM-4-9B-0414, Stage=10%2025.10 | 7.5 | — | — | |
| R-ZeroBackbone Model=OctoThinker-8B-Hybrid-Base2026.02 | 6.8 | — | — | |
| HACPOModel Backbone=Llama3.2-3B-Instruct2026.03 | 6.7 | — | — | |
| AlphaRLBase Model=Qwen3-8B-Base, Stage=5%2025.10 | 6.67 | — | — | |
| Base ModelBackbone Model=Qwen3-4B-Base2026.02 | 5.8 | — | — | |
| Standard TrainingBase Model=Qwen3-8B-Base, Stage=5%2025.10 | 5.5 | — | — | |
| Base ModelBackbone Model=OctoThinker-8B-Hybrid-Base2026.02 | 3.9 | — | — | |
| AlphaRLBase Model=GLM-4-9B-0414, Stage=5%2025.10 | 3.33 | — | — | |
| HACPOModel Backbone=Llama3.2-1B-Instruct, Experimental Context=Heterogeneous Setup2026.03 | 3.3 | — | — | |
| HACPOModel Backbone=Llama3.2-1B-Instruct2026.03 | 2.2 | — | — | |
| Standard TrainingBase Model=GLM-4-9B-0414, Stage=5%2025.10 | 1.67 | — | — |