Mathematical Reasoning on AIME 24 (Accuracy, Tokens)
93.3AccuracyDDC
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DDCBackbone=Qwen3-32B, Voting size budget=5122026.05 | 93.3 | 1.3 | — | |
| D-LBackbone=DeepSeek-R1-0528-Qwen3-8B, Voting size budget=5122026.05 | 92.1 | 7.8 | — | |
| D-LBackbone=Qwen3-32B, Voting size budget=5122026.05 | 89.4 | 7.8 | — | |
| SCBackbone=DeepSeek-R1-0528-Qwen3-8B, Voting size budget=5122026.05 | 86.7 | 35.5 | — | |
| D-HBackbone=DeepSeek-R1-0528-Qwen3-8B, Voting size budget=5122026.05 | 86.7 | 14.5 | — | |
| DDCBackbone=DeepSeek-R1-0528-Qwen3-8B, Voting size budget=5122026.05 | 86.7 | 1.9 | — | |
| ParaManagerBackbone=ParaManager-4B, Orchestration Strategy=Unified, Sampling Protocol (mean@8)=true2026.04 | 86.67 | — | — | |
| ACBackbone=DeepSeek-R1-0528-Qwen3-8B, Voting size budget=5122026.05 | 86.6 | 11.8 | — | |
| D-HBackbone=Qwen3-32B, Voting size budget=5122026.05 | 86.4 | 8.8 | — | |
| ACBackbone=Qwen3-32B, Voting size budget=5122026.05 | 84.9 | 10.6 | — | |
| SCBackbone=Qwen3-32B, Voting size budget=5122026.05 | 84.8 | 20 | — | |
| Meta Agent SearchBackbone=Qwen3-30B-A3B-Instruct-2507, Orchestration Strategy=Static Workflow, Sampling Protocol (mean@8)=true2026.04 | 84.17 | — | — | |
| EvoFlowOrchestration Strategy=Static Workflow, Sampling Protocol (mean@8)=true2026.04 | 84.17 | — | — | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Majority Vote, Sampling Protocol (mean@8)=true2026.04 | 83.33 | — | — | |
| GPT-OSS-20BOrchestration Strategy=Majority Vote, Sampling Protocol (mean@8)=true2026.04 | 83.33 | — | — | |
| GPT-OSS-20BOrchestration Strategy=Self Refine, Sampling Protocol (mean@8)=true2026.04 | 83.33 | — | — | |
| ParaManager-MonoBackbone=ParaManager-4B, Orchestration Strategy=Mono, Sampling Protocol (mean@8)=true2026.04 | 83.33 | — | — | |
| DDCBackbone=Qwen3-4B, Voting size budget=5122026.05 | 83.3 | 1.2 | — | |
| PathCalBase Model=QwQ-32B2026.05 | 83.3 | 13,426 | — | |
| D-LBackbone=Qwen3-4B, Voting size budget=5122026.05 | 83.1 | 14.1 | — | |
| D-LBackbone=Qwen3-8B, Voting size budget=5122026.05 | 80.4 | 9 | — | |
| LessIsMoreModel=Qwen3-14B, Token Budget=4K2025.08 | 80.39 | — | — | |
| D-HBackbone=Qwen3-4B, Voting size budget=5122026.05 | 80.2 | 24.2 | — | |
| LessIsMoreModel=Qwen3-14B, Token Budget=6K2025.08 | 80.19 | — | — | |
| LessIsMoreModel=Qwen3-14B, Token Budget=8K2025.08 | 80.1 | — | — | |
| D-HBackbone=Qwen3-8B, Voting size budget=5122026.05 | 80.1 | 13.3 | — | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Self Refine, Sampling Protocol (mean@8)=true2026.04 | 80 | — | — | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Majority Vote, Sampling Protocol (mean@8)=true2026.04 | 80 | — | — | |
| ToolOrchestraOrchestration Strategy=Serial Orchestration, Sampling Protocol (mean@8)=true2026.04 | 80 | — | — | |
| ParaManager-SFTBackbone=ParaManager-4B, Orchestration Strategy=SFT, Sampling Protocol (mean@8)=true2026.04 | 80 | — | — | |
| Qwen3-8B (Ours)Backbone=Qwen3-8B2026.04 | 80 | — | — | |
| SCBackbone=Qwen3-4B, Voting size budget=5122026.05 | 80 | 38.4 | — | |
| SCBackbone=Qwen3-8B, Voting size budget=5122026.05 | 80 | 23.2 | — | |
| ACBackbone=Qwen3-8B, Voting size budget=5122026.05 | 80 | 12.2 | — | |
| DDCBackbone=Qwen3-8B, Voting size budget=5122026.05 | 80 | 1.3 | — | |
| ACBackbone=Qwen3-4B, Voting size budget=5122026.05 | 79.9 | 12.8 | — | |
| Full AttnModel=Qwen3-14B2025.08 | 79.79 | — | — | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Self Refine, Sampling Protocol (mean@8)=true2026.04 | 79.17 | — | — | |
| LessIsMoreModel=Qwen3-14B, Token Budget=2K2025.08 | 78.58 | — | — | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Base, Sampling Protocol (mean@8)=true2026.04 | 77.5 | — | — | |
| RouterOrchestration Strategy=Static Workflow, Sampling Protocol (mean@8)=true2026.04 | 77.5 | — | — | |
| ParaManager-SerialBackbone=ParaManager-4B, Orchestration Strategy=Serial, Sampling Protocol (mean@8)=true2026.04 | 77.5 | — | — | |
| D-LBackbone=Qwen3-1.7B, Voting size budget=5122026.05 | 76.7 | 9.6 | — | |
| D-HBackbone=Qwen3-1.7B, Voting size budget=5122026.05 | 76.7 | 16.8 | — | |
| TIPBase Model=QwQ-32B2026.05 | 76.7 | 12,854 | — | |
| GPT-OSS-20BOrchestration Strategy=Base+Tool, Sampling Protocol (mean@8)=true2026.04 | 76.67 | — | — | |
| LessIsMoreModel=Qwen3-8B, Token Budget=8K2025.08 | 76.67 | — | — | |
| LessIsMoreModel=Qwen3-8B, Token Budget=6K2025.08 | 76.45 | — | — | |
| PuppeteerOrchestration Strategy=Serial Orchestration, Sampling Protocol (mean@8)=true2026.04 | 75.83 | — | — | |
| LessIsMoreModel=Qwen3-8B, Token Budget=4K2025.08 | 75.56 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 75.52 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 74.58 | — | — | |
| StepFlowBackbone=DeepSeek-R1-Distill-Qwen 32B2026.04 | 74.5 | — | — | |
| Full AttnModel=Qwen3-8B2025.08 | 74.48 | — | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=6K2025.08 | 74.37 | — | — | |
| QuestAHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 74.26 | — | — | |
| QuestAHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 74.23 | — | — | |
| SCBackbone=Qwen3-1.7B, Voting size budget=5122026.05 | 74.1 | 49 | — | |
| Plan-and-Solve (PS+)Backbone=DeepSeek-R1-Distill-Qwen 32B2026.04 | 74 | — | — | |
| Act. SteeringBackbone=DeepSeek-R1-Distill-Qwen 32B2026.04 | 74 | — | — | |
| ACBackbone=Qwen3-1.7B, Voting size budget=5122026.05 | 74 | 19.3 | — | |
| Attn-Interv.Backbone=DeepSeek-R1-Distill-Qwen 32B2026.04 | 73.5 | — | — | |
| GPT-OSS-20BOrchestration Strategy=Base, Sampling Protocol (mean@8)=true2026.04 | 73.33 | — | — | |
| MURBackbone=Qwen3-8B, Reasoning Mode=Thinking Mode2025.07 | 73.33 | 14,416 | — | |
| DDCBackbone=Qwen3-1.7B, Voting size budget=5122026.05 | 73.3 | 3.8 | — | |
| OriginalBase Model=QwQ-32B2026.05 | 73.3 | 12,886 | — | |
| S1Base Model=QwQ-32B2026.05 | 73.3 | 14,575 | — | |
| CyclicReflexBase Model=QwQ-32B2026.05 | 73.3 | 12,848 | — | |
| Budget Forcing (S1)Backbone=DeepSeek-R1-Distill-Qwen 32B2026.04 | 73.2 | — | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=8K2025.08 | 73.12 | — | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=4K2025.08 | 73.03 | — | — | |
| Hint-Infer (Round1)Backbone=DeepSeek-R1-Distill-Qwen 32B2026.04 | 73 | — | — | |
| LessIsMoreModel=Qwen3-8B, Token Budget=2K2025.08 | 73 | — | — | |
| DeepSeek-R1-Distill-Qwen 32BInference Protocol=Baseline2026.04 | 72.6 | — | — | |
| Per-Step ScaleBackbone=Qwen3-8B, Reasoning Mode=Thinking Mode2025.07 | 72.29 | 13,793 | — | |
| StepFlowBackbone=DeepSeek-R1-Distill-Qwen 14B2026.04 | 72.1 | — | — | |
| Plan-and-Solve (PS+)Backbone=DeepSeek-R1-Distill-Qwen 14B2026.04 | 71.8 | — | — | |
| QuestAHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 71.56 | — | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=2K2025.08 | 71.48 | — | — | |
| Full AttnModel=Qwen3-4B2025.08 | 71.25 | — | — | |
| Hint-Infer (Round1)Backbone=DeepSeek-R1-Distill-Qwen 14B2026.04 | 71.2 | — | — | |
| Act. SteeringBackbone=DeepSeek-R1-Distill-Qwen 14B2026.04 | 71 | — | — | |
| Attn-Interv.Backbone=DeepSeek-R1-Distill-Qwen 14B2026.04 | 70.8 | — | — | |
| Budget Forcing (S1)Backbone=DeepSeek-R1-Distill-Qwen 14B2026.04 | 70.5 | — | — | |
| JustRLHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 70.42 | — | — | |
| Avg uncertaintyBackbone=Qwen3-8B, Reasoning Mode=Thinking Mode2025.07 | 70.42 | 15,463 | — | |
| Qwen3-4B (Ours)Backbone=Qwen3-4B2026.04 | 70 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 69.79 | — | — | |
| JustRLHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 69.76 | — | — | |
| DeepSeek-R1-Distill-Qwen 14BInference Protocol=Baseline2026.04 | 69.7 | — | — | |
| JustRLHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 69.69 | — | — | |
| Avg uncertaintyBackbone=Qwen3-4B, Reasoning Mode=Thinking Mode2025.07 | 68.75 | 14,832 | — | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Base, Sampling Protocol (mean@8)=true2026.04 | 68.33 | — | — | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Base+Tool, Sampling Protocol (mean@8)=true2026.04 | 68.33 | — | — | |
| Per-Step ScaleBackbone=Qwen3-4B, Reasoning Mode=Thinking Mode2025.07 | 68.33 | 13,648 | — | |
| SMARTBackbone=Qwen3-4B, Reasoning Mode=Thinking Mode2025.07 | 68.33 | 15,131 | — | |
| SMARTBackbone=Qwen3-8B, Reasoning Mode=Thinking Mode2025.07 | 68.33 | 16,926 | — | |
| MURBackbone=Qwen3-4B, Reasoning Mode=Thinking Mode2025.07 | 68.13 | 13,009 | — | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Base+Tool, Sampling Protocol (mean@8)=true2026.04 | 66.67 | — | — | |
| StepFlowBackbone=GPT-OSS-20B medium2026.04 | 66 | — | — |