Mathematical Reasoning on AIME 24 (test)
97.5AccuracyIterative Critique-and-Routing Controller
Evaluation Results
| Method | Links | |
|---|---|---|
| Iterative Critique-and-Routing ControllerController Backbone=Qwen3-8B-Base2026.05 | 97.5 | |
| Agent 1Controller Backbone=Qwen3-4B-Base, Agent 1 Model=Qwen3-30B-A3B-Instruct2026.05 | 83.3 | |
| Iterative Critique-and-Routing ControllerController Backbone=Qwen3-4B-Base2026.05 | 83.3 | |
| Agent 1Controller Backbone=Qwen3-8B-Base, Agent 1 Model=Qwen3-30B-A3B-Instruct2026.05 | 83.3 | |
| RoBERTa RouterController Backbone=Qwen3-8B-Base2026.05 | 83.3 | |
| Router-R1Controller Backbone=Qwen3-8B-Base2026.05 | 80 | |
| Router-R1Controller Backbone=Qwen3-4B-Base2026.05 | 76.7 | |
| RoBERTa RouterController Backbone=Qwen3-4B-Base2026.05 | 66.7 | |
| Solver + Verifier (AstraFlow)Backbone=Qwen3-8B, Framework=AstraFlow, Time/iter (s)=77.652026.05 | 47.3 | |
| RouterDCController Backbone=Qwen3-8B-Base2026.05 | 46.7 | |
| Solver + Verifier (verl)Backbone=Qwen3-8B, Framework=verl, Time/iter (s)=212.642026.05 | 44.6 | |
| SolverBackbone=Qwen3-8B2026.05 | 42.9 | |
| Controller V2Controller Backbone=Qwen3-4B-Base2026.05 | 33.3 | |
| Random RouterController Backbone=Qwen3-4B-Base2026.05 | 27.5 | |
| RouterDCController Backbone=Qwen3-4B-Base2026.05 | 27.5 | |
| Random RouterController Backbone=Qwen3-8B-Base2026.05 | 26.7 | |
| GAC + Token-φVariant=Token-φ2026.05 | 20.8 | |
| Controller V2Controller Backbone=Qwen3-8B-Base2026.05 | 20 | |
| GAC w/o φVariant=w/o φ2026.05 | 20 | |
| HPTCategory=Recent hybrid SFT–RL2026.05 | 18.7 | |
| LUFFYCategory=Recent hybrid SFT–RL2026.05 | 18.5 | |
| KL-ctrlCategory=Rule-based controllers2026.05 | 18.4 | |
| CHORDCategory=SFT–RL mixing baselines2026.05 | 18.2 | |
| Nash-MTLCategory=Multi-objective solvers2026.05 | 18.1 | |
| CAGradCategory=Multi-objective solvers2026.05 | 18 | |
| GradNorm-ctrlCategory=Rule-based controllers2026.05 | 18 | |
| SRFTCategory=Recent hybrid SFT–RL2026.05 | 17.9 | |
| SFT-best + RLCategory=SFT–RL mixing baselines2026.05 | 17.1 | |
| Controller V1Controller Backbone=Qwen3-8B-Base2026.05 | 16.7 | |
| DPOCategory=RL-free alignment2026.05 | 16.4 | |
| IPOCategory=RL-free alignment2026.05 | 16.1 | |
| SFT-best2026.05 | 15.8 | |
| Controller V1Controller Backbone=Qwen3-4B-Base2026.05 | 13.3 | |
| GRPO (pure RL)Category=SFT–RL mixing baselines2026.05 | 13.2 | |
| Qwen2.5-7B-Inst.2026.05 | 11.7 | |
| Agent 2Controller Backbone=Qwen3-4B-Base, Agent 2 Model=Qwen2.5-7B-Instruct2026.05 | 10 | |
| Agent 2Controller Backbone=Qwen3-8B-Base, Agent 2 Model=Ministral-3-8B-Instruct2026.05 | 10 | |
| BaseBackbone=LLaMA 3.1-8B-Instruct, Optimization Setting=test-time optimization, Reward Supervision=noisy rule-based2025.08 | 10 | |
| VRPOBackbone=LLaMA 3.1-8B-Instruct, Optimization Setting=test-time optimization, Reward Supervision=noisy rule-based2025.08 | 10 | |
| Agent 3Controller Backbone=Qwen3-4B-Base, Agent 3 Model=Qwen2.5-1.5B-Instruct2026.05 | 6.7 | |
| GRPOBackbone=LLaMA 3.1-8B-Instruct, Optimization Setting=test-time optimization, Reward Supervision=noisy rule-based2025.08 | 6.67 | |
| PPOBackbone=LLaMA 3.1-8B-Instruct, Optimization Setting=test-time optimization, Reward Supervision=noisy rule-based2025.08 | 6.67 | |
| Agent 3Controller Backbone=Qwen3-8B-Base, Agent 3 Model=Llama-3-2-1B-Instruct2026.05 | 0 |