Mathematical Reasoning on AMC 23 (Acc, Len, CR)
98.8AccuracyVanilla
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VanillaBackbone=Gemini2.5-Flash2026.04 | 98.8 | 2,290 | 100 | |
| TRACEBackbone=Gemini2.5-Flash2026.04 | 98.5 | 1,994 | 87.4 | |
| Agent 1 (Qwen3-30B-A3B-Instruct)Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 97.5 | — | — | |
| Agent 1 (Qwen3-30B-A3B-Instruct)Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 97.5 | — | — | |
| Iterative Critique-and-Routing ControllerController=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 97.5 | — | — | |
| VanillaBackbone=Qwen3-8B2026.04 | 96.6 | 7,885 | 100 | |
| NoThinkBackbone=Gemini2.5-Flash2026.04 | 96.5 | 2,311 | 101 | |
| JustRLHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 96.02 | — | — | |
| TRACEBackbone=Qwen3-8B2026.04 | 96 | 5,405 | 68.6 | |
| DynasorBackbone=Qwen3-8B2026.04 | 95.9 | 6,552 | 83.1 | |
| KnowRL-Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 95.78 | — | — | |
| JustRLHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 95.7 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 95.7 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 95.55 | — | — | |
| JustRLHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 95.54 | — | — | |
| DynasorBackbone=Gemini2.5-Flash2026.04 | 95.3 | 2,001 | 87.4 | |
| QuestAHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 95.1 | — | — | |
| QuestAHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 95.08 | — | — | |
| DEERBackbone=Qwen3-8B2026.04 | 94.7 | 6,455 | 81.9 | |
| TALEBackbone=Qwen3-8B2026.04 | 94.1 | 5,890 | 74.7 | |
| VanillaBackbone=Qwen3-4B2026.04 | 93.7 | 8,089 | 100 | |
| DEERBackbone=Qwen3-4B2026.04 | 93.5 | 6,588 | 81.5 | |
| QuestAHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 93.44 | — | — | |
| TALEBackbone=Qwen3-4B2026.04 | 93.1 | 6,772 | 83.7 | |
| DEERBackbone=Gemini2.5-Flash2026.04 | 93.1 | 1,874 | 81.8 | |
| TRACEBackbone=Qwen3-4B2026.04 | 92.7 | 5,652 | 69.9 | |
| O1-PrunerBackbone=DeepSeek-R1-Distill-7B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 92.5 | 3.2 | 48.4 | |
| PEARBackbone=DeepSeek-R1-Distill-7B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 92.5 | 4.4 | 66.04 | |
| DEERBackbone=Qwen3-4B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 92.5 | 6.1 | 80.3 | |
| PEARBackbone=Qwen3-4B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 92.5 | 6 | 34.2 | |
| DEERBackbone=Qwen3-8B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 92.5 | 6.4 | 80 | |
| ETRBackbone=Qwen3-8B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 92.5 | 4.2 | 53.8 | |
| Iterative Critique-and-Routing ControllerController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 92.5 | — | — | |
| Router-R1Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 92.5 | — | — | |
| TALEBackbone=Gemini2.5-Flash2026.04 | 92.3 | 1,856 | 81 | |
| VanillaBackbone=R1-Distilled-Llama-8B2026.04 | 91.7 | 5,717 | 100 | |
| Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 91.56 | — | — | |
| Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 90.7 | — | — | |
| Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 90.47 | — | — | |
| DynasorBackbone=Qwen3-4B2026.04 | 90.1 | 6,812 | 84.2 | |
| OriginalBackbone=Qwen3-4B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 90 | 7.6 | 100 | |
| LCPOBackbone=Qwen3-4B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 90 | 6 | 78.9 | |
| ETRBackbone=Qwen3-4B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 90 | 4 | 52.6 | |
| OriginalBackbone=Qwen3-8B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 90 | 8 | 100 | |
| LCPOBackbone=Qwen3-8B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 90 | 5.7 | 41.3 | |
| O1-PrunerBackbone=Qwen3-8B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 90 | 3.3 | 41.3 | |
| DynasorBackbone=R1-Distilled-Llama-8B2026.04 | 89.3 | 4,811 | 84.2 | |
| TRACEBackbone=R1-Distilled-Llama-8B2026.04 | 88.5 | 4,576 | 80.1 | |
| NoThinkBackbone=R1-Distilled-Llama-8B2026.04 | 88.1 | 5,288 | 92.5 | |
| LCPOBackbone=DeepSeek-R1-Distill-7B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 87.5 | 3.5 | 53 | |
| ETRBackbone=DeepSeek-R1-Distill-7B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 87.5 | 2.4 | 36.4 | |
| PEARBackbone=Qwen3-8B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 87.5 | 6.8 | 84.7 | |
| DEERBackbone=DeepSeek-R1-Distill-7B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 85 | 4.9 | 74.2 | |
| O1-PrunerBackbone=Qwen3-4B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 85 | 2.6 | 34.2 | |
| DEERBackbone=R1-Distilled-Llama-8B2026.04 | 84.1 | 4,551 | 79.6 | |
| EfficientReasoningModel Scale=7B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 83.44 | 3,401 | — | |
| NoThinkBackbone=Qwen3-4B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 82.5 | 2.4 | 31.6 | |
| TALEBackbone=R1-Distilled-Llama-8B2026.04 | 82.5 | 3,169 | 55.4 | |
| Router-R1Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 82.5 | — | — | |
| Graph-Based Chain-of-Thought PruningModel Scale=7B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 82.34 | 3,211 | — | |
| OREAL-7BModel Scale=7B, Method Variant=Open-Source R1-Style2026.04 | 81.25 | 4,510 | — | |
| BaseModel Scale=7B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 81 | 6,849 | — | |
| AdaptThinkModel Scale=7B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 80.75 | 3,531 | — | |
| OriginalBackbone=DeepSeek-R1-Distill-7B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 80 | 6.6 | 100 | |
| TokenSkipModel Scale=7B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 79.06 | 4,811 | — | |
| Light-R1-DS-7BModel Scale=7B, Method Variant=Open-Source R1-Style2026.04 | 78.25 | 6,073 | — | |
| NoThinkBackbone=DeepSeek-R1-Distill-7B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 77.5 | 3.8 | 57.6 | |
| Controller V2Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 77.5 | — | — | |
| RoBERTa RouterController=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 77.5 | — | — | |
| O1-PrunerModel Scale=7B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 75.62 | 5,128 | — | |
| Fixed-MAS + ADv1Framework=Fixed-MAS, Technique=ADv12026.02 | 75 | — | — | |
| NoThinkBackbone=Qwen3-8B2026.04 | 74.8 | 1,575 | 20 | |
| Skywork-OR1-7BModel Scale=7B, Method Variant=Open-Source R1-Style2026.04 | 73.5 | 5,241 | — | |
| NoThinkBackbone=Qwen3-4B2026.04 | 72.5 | 1,814 | 22.4 | |
| AReaL-boba-RL-7BModel Scale=7B, Method Variant=Open-Source R1-Style2026.04 | 71.56 | 5,337 | — | |
| Fixed-MAS + PRMFramework=Fixed-MAS, Technique=PRM2026.02 | 70 | — | — | |
| Fixed-MAS + Multi-TAGFramework=Fixed-MAS, Technique=Multi-TAG2026.02 | 70 | — | — | |
| Dynamic-MAS + ADv2Framework=Dynamic-MAS, Technique=ADv22026.02 | 70 | — | — | |
| Graph-Based Chain-of-Thought PruningModel Scale=1.5B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 69.38 | 2,774 | — | |
| EfficientReasoningModel Scale=1.5B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 68.75 | 3,475 | — | |
| NoThinkBackbone=Qwen3-8B, Decoding Strategy=greedy decoding (pass@1)2026.04 | 67.5 | 3.4 | 41.3 | |
| Single Agent + CoTFramework=Single Agent, Technique=Chain-of-Thought2026.02 | 67.5 | — | — | |
| Fixed-MAS + ADv2Framework=Fixed-MAS, Technique=ADv22026.02 | 67.5 | — | — | |
| Dynamic-MAS + PRMFramework=Dynamic-MAS, Technique=PRM2026.02 | 67.5 | — | — | |
| AdaptThinkModel Scale=1.5B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 65.31 | 2,039 | — | |
| Fixed-MASFramework=Fixed-MAS2026.02 | 65 | — | — | |
| Fixed-MAS + Self-RefineFramework=Fixed-MAS, Technique=Self-Refine2026.02 | 65 | — | — | |
| BaseModel Scale=1.5B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 63.12 | 5,205 | — | |
| Single AgentFramework=Single Agent2026.02 | 62.5 | — | — | |
| Dynamic-MASFramework=Dynamic-MAS2026.02 | 62.5 | — | — | |
| Dynamic-MAS + Self-RefineFramework=Dynamic-MAS, Technique=Self-Refine2026.02 | 62.5 | — | — | |
| Dynamic-MAS + Multi-TAGFramework=Dynamic-MAS, Technique=Multi-TAG2026.02 | 62.5 | — | — | |
| TokenSkipModel Scale=1.5B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 60.62 | 5,452 | — | |
| Controller V2Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 60 | — | — | |
| RouterDCController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 60 | — | — | |
| RouterDCController=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 60 | — | — | |
| O1-PrunerModel Scale=1.5B, Backbone=DeepSeek-R1-Distill-Qwen2026.04 | 59.38 | 5,798 | — | |
| Random RouterController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 57.5 | — | — | |
| RoBERTa RouterController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 55 | — | — | |
| Controller V1Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 55 | — | — |