Mathematical Reasoning on AMC (Accuracy %)
95.18Accuracy (%)Qwen3-30B A3B-Instruct-2507
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Majority Vote, Sampling Protocol (mean@8)=true2026.04 | 95.18 | |
| ParaManagerBackbone=ParaManager-4B, Orchestration Strategy=Unified, Sampling Protocol (mean@8)=true2026.04 | 95.18 | |
| Meta Agent SearchOrchestration Strategy=Static Workflow, Sampling Protocol (mean@8)=true2026.04 | 94.58 | |
| GPT-OSS-20BOrchestration Strategy=Base+Tool, Sampling Protocol (mean@8)=true2026.04 | 93.98 | |
| PuppeteerOrchestration Strategy=Serial Orchestration, Sampling Protocol (mean@8)=true2026.04 | 93.98 | |
| EvoFlowOrchestration Strategy=Static Workflow, Sampling Protocol (mean@8)=true2026.04 | 92.47 | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Self Refine, Sampling Protocol (mean@8)=true2026.04 | 92.17 | |
| GPT-OSS-20BOrchestration Strategy=Majority Vote, Sampling Protocol (mean@8)=true2026.04 | 92.17 | |
| GPT-OSS-20BOrchestration Strategy=Self Refine, Sampling Protocol (mean@8)=true2026.04 | 92.17 | |
| ParaManager-MonoBackbone=ParaManager-4B, Orchestration Strategy=Mono, Sampling Protocol (mean@8)=true2026.04 | 92.17 | |
| ParaManager-SFTBackbone=ParaManager-4B, Orchestration Strategy=SFT, Sampling Protocol (mean@8)=true2026.04 | 92.02 | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Base, Sampling Protocol (mean@8)=true2026.04 | 91.87 | |
| RouterOrchestration Strategy=Static Workflow, Sampling Protocol (mean@8)=true2026.04 | 91.87 | |
| ParaManager-SerialBackbone=ParaManager-4B, Orchestration Strategy=Serial, Sampling Protocol (mean@8)=true2026.04 | 91.87 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Majority Vote, Sampling Protocol (mean@8)=true2026.04 | 91.57 | |
| GPT-OSS-20BOrchestration Strategy=Base, Sampling Protocol (mean@8)=true2026.04 | 90.66 | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Base+Tool, Sampling Protocol (mean@8)=true2026.04 | 90.36 | |
| ToolOrchestraOrchestration Strategy=Serial Orchestration, Sampling Protocol (mean@8)=true2026.04 | 90.36 | |
| HAPOBackbone=Qwen3-14B2025.09 | 89.66 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Self Refine, Sampling Protocol (mean@8)=true2026.04 | 89.61 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Base+Tool, Sampling Protocol (mean@8)=true2026.04 | 88.25 | |
| DAPO w/ Forking TokensBackbone=Qwen3-14B2025.09 | 87.31 | |
| EDGE-GRPOBackbone=Qwen3-14B2025.09 | 87.15 | |
| Entropy EnvBackbone=Qwen3-14B2025.09 | 87.04 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Base, Sampling Protocol (mean@8)=true2026.04 | 86.9 | |
| ArcherBackbone=Qwen3-14B2025.09 | 86.73 | |
| Vanilla DAPOBackbone=Qwen3-14B2025.09 | 85.52 | |
| HAPOBackbone=Qwen2.5-Math-7B2025.09 | 85.47 | |
| Vanilla GRPOBackbone=Qwen3-14B2025.09 | 83.91 | |
| EDGE-GRPOBackbone=Qwen2.5-Math-7B2025.09 | 83.23 | |
| Entropy EnvBackbone=Qwen2.5-Math-7B2025.09 | 82.94 | |
| ArcherBackbone=Qwen2.5-Math-7B2025.09 | 82.53 | |
| Vanilla DAPOBackbone=Qwen2.5-Math-7B2025.09 | 81.99 | |
| HAPOBackbone=Qwen3-8B2025.09 | 81.77 | |
| DAPO w/ Forking TokensBackbone=Qwen2.5-Math-7B2025.09 | 80.55 | |
| EDGE-GRPOBackbone=Qwen3-8B2025.09 | 80.47 | |
| DAPO w/ Forking TokensBackbone=Qwen3-8B2025.09 | 80.26 | |
| Vanilla GRPOBackbone=Qwen2.5-Math-7B2025.09 | 80.14 | |
| Entropy EnvBackbone=Qwen3-8B2025.09 | 80.05 | |
| ArcherBackbone=Qwen3-8B2025.09 | 78.47 | |
| Vanilla DAPOBackbone=Qwen3-8B2025.09 | 78.37 | |
| TAPOBase Model=Qwen2.5-Math-7B2025.07 | 77.5 | |
| TemplateRLBackbone=Qwen2.5-Math-7B-Base, RL Baseline=GRPO2025.05 | 77.5 | |
| Vanilla GRPOBackbone=Qwen3-8B2025.09 | 76.84 | |
| HAPOBackbone=Qwen2.5-Math-1.5B2025.09 | 75.69 | |
| AISModel=Qwen3.5-9B, Bitwidth=FP82026.05 | 75.5 | |
| DAPO w/ Forking TokensBackbone=Qwen2.5-Math-1.5B2025.09 | 74.8 | |
| EDGE-GRPOBackbone=Qwen2.5-Math-1.5B2025.09 | 74.48 | |
| RL (BF16)Model=Qwen3.5-9B, Bitwidth=BF162026.05 | 74.2 | |
| Entropy EnvBackbone=Qwen2.5-Math-1.5B2025.09 | 72.88 | |
| Vanilla DAPOBackbone=Qwen2.5-Math-1.5B2025.09 | 72.85 | |
| ArcherBackbone=Qwen2.5-Math-1.5B2025.09 | 72.17 | |
| Vanilla GRPOBackbone=Qwen2.5-Math-1.5B2025.09 | 71.23 | |
| SPICEBackbone=Qwen3-8B-Base, #Data=20K2026.05 | 70 | |
| AISModel=Qwen3-8B, Bitwidth=FP82026.05 | 69.75 | |
| DDRLBase Model=Qwen2.5-Math-7B2026.04 | 69 | |
| RL (BF16)Model=Qwen3-8B, Bitwidth=BF162026.05 | 68.5 | |
| RL-PLUSBase Model=Qwen2.5-Math-7B, Training Method=RL-PLUS2025.07 | 68.1 | |
| RL-PLUSBase Model=Qwen2.5-Math-7B2025.07 | 68.1 | |
| TTRLBase Model=Qwen2.5-Math-7B2026.04 | 68.1 | |
| FlashRL (TIS)Model=Qwen3.5-9B, Bitwidth=FP82026.05 | 67.5 | |
| TemplateRLBackbone=Qwen2.5-Math-7B-Base, RL Baseline=RLOO2025.05 | 67.5 | |
| TEPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.04 | 66.41 | |
| Hi-CoTModel=Qwen3-4B-Instruct-25072026.03 | 66.3 | |
| DAPOBase Model=Qwen2.5-Math-7B2025.07 | 66.3 | |
| GRPO/DAPOBackbone=DeepSeek-R1-Distill-Qwen-7B2026.04 | 66.1 | |
| DoTS (ExGRPO + ReLIFT)Backbone=Qwen2.5-Math-7B, Integration=Decoupled Test-time Synthesis2026.05 | 66.1 | |
| RL (FP8 Rollout)Model=Qwen3.5-9B, Bitwidth=FP82026.05 | 65.8 | |
| ExGRPOBackbone=Qwen2.5-Math-7B2026.05 | 65.7 | |
| LUFFYBase Model=Qwen2.5-Math-7B2025.07 | 65.6 | |
| CoTModel=Qwen3-4B-Instruct-25072026.03 | 65.1 | |
| Plan-and-SolveModel=Qwen3-4B-Instruct-25072026.03 | 65.1 | |
| PivotTraceBackbone=Deepseek-R1-Distill-Qwen-1.5B2026.06 | 65.1 | |
| QueSTBackbone=Qwen3-8B-Base2026.05 | 65 | |
| CoEBackbone=Deepseek-R1-Distill-Qwen-1.5B2026.06 | 65 | |
| CoT-KineticsBackbone=Deepseek-R1-Distill-Qwen-1.5B2026.06 | 64.8 | |
| EntropyBackbone=Deepseek-R1-Distill-Qwen-1.5B2026.06 | 64.8 | |
| D2EvoBackbone=Qwen3-8B-Base, #Data=0.3K, Iteration=Iter 22026.05 | 64.76 | |
| D2EvoBackbone=Qwen3-8B-Base, #Data=0.4K, Iteration=Iter 32026.05 | 64.76 | |
| LUFFYBackbone=Qwen2.5-Math-7B2026.05 | 64.6 | |
| ReLIFTBackbone=Qwen2.5-Math-7B2026.05 | 64.4 | |
| D2EvoBackbone=Qwen3-4B-Base, #Data=0.1K, Iteration=Iter 32026.05 | 64.38 | |
| D2EvoBackbone=Qwen3-8B-Base, #Data=1K, Iteration=Iter 12026.05 | 64.34 | |
| ReLIFTBase Model=Qwen2.5-Math-7B2025.07 | 64.3 | |
| FlashRL (TIS)Model=Qwen3-8B, Bitwidth=FP82026.05 | 63.4 | |
| ConsistencyBackbone=Deepseek-R1-Distill-Qwen-1.5B, Multiple stochastic inferences=true2026.06 | 63.4 | |
| DoTS (SFT + On-Policy RL)Backbone=Qwen2.5-Math-7B, Integration=Decoupled Test-time Synthesis2026.05 | 63.2 | |
| Full DataBackbone=Qwen3-8B-Base, #Data=19K2026.05 | 63.17 | |
| StandardModel=Qwen3-4B-Instruct-25072026.03 | 62.7 | |
| SFT+GRPOBase Model=Qwen2.5-Math-7B2025.07 | 62.7 | |
| SFT+RLBackbone=Qwen2.5-Math-7B, Protocol=Sequential training2026.05 | 62.7 | |
| Self-CertaintyBackbone=Deepseek-R1-Distill-Qwen-1.5B2026.06 | 62.6 | |
| TLMBackbone=Qwen3-8B-Base2026.05 | 62.5 | |
| AZRBackbone=Qwen3-8B-Base2026.05 | 62.5 | |
| Full DataBackbone=Qwen3-4B-Base, #Data=19K2026.05 | 62.42 | |
| TIES-Merging* (SFT + On-Policy RL)Backbone=Qwen2.5-Math-7B, Integration=Model Merging with DoTS sparsification2026.05 | 62.3 | |
| GRPOBackbone=Qwen2.5-Math-7B2026.04 | 62 | |
| GRPOBase Model=Qwen2.5-Math-7B, Training Method=GRPO2025.07 | 62 | |
| GRPOBase Model=Qwen2.5-Math-7B2025.07 | 62 | |
| On-Policy RLBackbone=Qwen2.5-Math-7B2026.05 | 62 |