Mathematical Reasoning on AQUA
85.05AccuracyOFA-MAS (fine-tuned)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OFA-MAS (fine-tuned)Training protocol=One-for-All Topology Design (Fine-tuned)2026.01 | 85.05 | — | |
| EIB-LEARNERTraining protocol=Dataset-Specific Training2026.01 | 83.49 | — | |
| OFA-MAS (pre-trained)Training protocol=One-for-All Topology Design (Pre-trained)2026.01 | 83.18 | — | |
| RAPSNumber of agents=52026.02 | 82.6 | — | |
| G-designerTraining protocol=Dataset-Specific Training2026.01 | 81.6 | — | |
| Teaching-Inspired Integrated Prompting FrameworkModel=GPT-42024.10 | 81.1 | — | |
| AgentDropoutTraining protocol=Dataset-Specific Training2026.01 | 80.94 | — | |
| SoftCoTN (Number of reasoning chains)=10, Base Model=Qwen3-8B2025.02 | 80.63 | — | |
| AgentPruneTraining protocol=Dataset-Specific Training2026.01 | 80.51 | — | |
| Previous SoTA2024.10 | 79.9 | — | |
| GPT-4 + PHPPHP=true, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 79.9 | 2.2913 | |
| G-DesignerNumber of agents=52026.02 | 79.4 | — | |
| IoTModel=GPT-4o mini2026.03 | 79.13 | — | |
| AgentPruneNumber of agents=52026.02 | 79.1 | — | |
| AFlowNumber of agents=52026.02 | 78.5 | — | |
| GPTSwarmNumber of agents=52026.02 | 78.2 | — | |
| LLM-DebateAgent topology=Fixed Multi-Agent (LLM-Debate)2026.01 | 77.65 | — | |
| GPT-4PHP=false, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 77.5 | — | |
| PuppeteerNumber of agents=52026.02 | 77.5 | — | |
| LLM-DebateNumber of agents=52026.02 | 77.3 | — | |
| GPT-42024.05 | 76.9 | — | |
| LLM-BlenderNumber of agents=52026.02 | 76.9 | — | |
| SCNumber of agents=12026.02 | 76.8 | — | |
| Zero-Shot CoTN (Number of reasoning chains)=10, Base Model=Qwen3-8B2025.02 | 76.77 | — | |
| Zero-Shot Assist-CoTN (Number of reasoning chains)=10, Base Model=Qwen3-8B2025.02 | 76.77 | — | |
| RandomAgent topology=Fixed Multi-Agent (Random)2026.01 | 76.48 | — | |
| CoTModel=GPT-42024.10 | 76.4 | — | |
| Previous SOTAPHP=false, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 76.4 | — | |
| CoconutN (Number of reasoning chains)=10, Base Model=Qwen3-8B2025.02 | 76.38 | — | |
| MaASNumber of agents=52026.02 | 76.2 | — | |
| ComplexCoTNumber of agents=12026.02 | 76.1 | — | |
| AutoAgentsNumber of agents=52026.02 | 75.7 | — | |
| SC (CoT)Prompting strategy=Single-Agent Prompting2026.01 | 75.63 | — | |
| RandomNumber of agents=52026.02 | 75.1 | — | |
| SoftCoTN (Number of reasoning chains)=1, Base Model=Qwen3-8B2025.02 | 75.04 | — | |
| AbstRaLModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=AbstRaL, Evaluation Protocol=Zero-shot2025.06 | 74.8 | — | |
| D-RPCStudent Model=Qwen 3 1.7B2026.05 | 74.76 | — | |
| CoTNumber of agents=12026.02 | 74.7 | — | |
| Ori-SFTModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=Ori-SFT, Evaluation Protocol=Zero-shot2025.06 | 74.4 | — | |
| CoTModel=GPT-4o mini2026.03 | 74.37 | — | |
| ChainAgent topology=Fixed Multi-Agent (Chain)2026.01 | 74.05 | — | |
| TreeNumber of agents=52026.02 | 73.9 | — | |
| SCModel=Olmo-2-13B2026.03 | 73.62 | — | |
| CoTPrompting strategy=Single-Agent Prompting2026.01 | 73.58 | — | |
| CompleteAgent topology=Fixed Multi-Agent (Complete)2026.01 | 72.95 | — | |
| MAS-ZeroNumber of agents=52026.02 | 72.9 | — | |
| CoAModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=CoA, Evaluation Protocol=Zero-shot2025.06 | 72.8 | — | |
| Vanilla IONumber of agents=1, Backbone=GPT-4o-mini2026.02 | 71.3 | — | |
| TreeAgent topology=Fixed Multi-Agent (Tree)2026.01 | 71.23 | — | |
| VanillaPrompting strategy=Direct Answer2026.01 | 71.06 | — | |
| Teaching-Inspired Integrated Prompting FrameworkModel=GPT-3.5-Turbo2024.10 | 70.8 | — | |
| IoTModel=Olmo-2-13B2026.03 | 70.47 | — | |
| ChainNumber of agents=52026.02 | 70.4 | — | |
| Zero-Shot Assist-CoTN (Number of reasoning chains)=1, Base Model=Qwen3-8B2025.02 | 70.16 | — | |
| Zero-Shot CoTN (Number of reasoning chains)=1, Base Model=Qwen3-8B2025.02 | 70 | — | |
| StarNumber of agents=52026.02 | 69.6 | — | |
| CoconutN (Number of reasoning chains)=1, Base Model=Qwen3-8B2025.02 | 68.5 | — | |
| D-RPCStudent Model=Llama 3.1 8B Instruct2026.05 | 67.52 | — | |
| CoTModel=Olmo-2-13B2026.03 | 67.32 | — | |
| SCModel=Olmo-2-7B2026.03 | 66.54 | — | |
| JiuZhang3.0Parameters=8x7B2024.05 | 65.4 | — | |
| CoT-RLModel=Qwen2.5-Math-7B-Instruct, GSM Train Method=CoT-RL, Evaluation Protocol=Zero-shot2025.06 | 65.4 | — | |
| Qwen-1.5Parameters=110B2024.05 | 64.6 | — | |
| CoTStudent Model=Llama 3.1 8B Instruct2026.05 | 64.02 | — | |
| IoTModel=Olmo-2-7B2026.03 | 63.78 | — | |
| SuperCorrectStudent Model=Qwen 3 1.7B2026.05 | 63.78 | — | |
| CoTModel=Olmo-2-7B2026.03 | 62.99 | — | |
| FreeformStudent Model=Qwen 3 1.7B2026.05 | 62.99 | — | |
| MAmmoTH2Parameters=7B, Variant=Plus2024.05 | 62.2 | — | |
| JiuZhang3.0Parameters=8B2024.05 | 62.2 | — | |
| IoTModel=Llama-3.3-8B2026.03 | 61.81 | — | |
| DCoTStudent Model=Llama 3.1 8B Instruct2026.05 | 61.81 | — | |
| DeepSeekMathParameters=7B, Variant=Instruct2024.05 | 60.6 | — | |
| CoTModel=GPT-3.5-Turbo2024.10 | 60.6 | — | |
| GPT-3.5 Turbo + PHPPHP=true, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 60.6 | 2.3228 | |
| FreeformStudent Model=Llama 3.1 8B Instruct2026.05 | 60.39 | — | |
| SGFTStudent Model=Llama 3.1 8B Instruct2026.05 | 60.16 | — | |
| CoTStudent Model=Qwen 3 1.7B2026.05 | 59.92 | — | |
| DCoTStudent Model=Qwen 3 1.7B2026.05 | 59.65 | — | |
| SuperCorrectStudent Model=Llama 3.1 8B Instruct2026.05 | 59.45 | — | |
| JiuZhang3.0Parameters=7B2024.05 | 59.4 | — | |
| Question-Analysis PromptingModel=GPT-3.5 Turbo, Word Count Constraint (n)=1502024.07 | 59.4 | — | |
| EoTModel=Olmo-2-13B2026.03 | 58.23 | — | |
| MAmmoTH2Parameters=8B, Variant=Plus2024.05 | 57.5 | — | |
| GPT-3.5 TurboPHP=false, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 57.4 | — | |
| Take A Deep BreathModel=GPT-3.5 Turbo, Prompt=TADB2024.07 | 57.1 | — | |
| SCModel=Llama-3.3-8B2026.03 | 57.09 | — | |
| SoftCoTBackbone=LLaMA-3.1-8B-Instruct2025.02 | 56.3 | — | |
| MAmmoTH2Parameters=8x7B, Variant=Plus2024.05 | 55.9 | — | |
| Zero-Shot Assist-CoTBackbone=LLaMA-3.1-8B-Instruct2025.02 | 55.83 | — | |
| EoTModel=Olmo-2-7B2026.03 | 55.82 | — | |
| Zero-Shot CoT-UnkBackbone=LLaMA-3.1-8B-Instruct2025.02 | 55.28 | — | |
| Qwen-1.5Parameters=72B2024.05 | 55.1 | — | |
| CoTModel=Llama-3.3-8B2026.03 | 54.72 | — | |
| Zero-Shot CoTBackbone=LLaMA-3.1-8B-Instruct2025.02 | 54.65 | — | |
| ChatGPT2024.05 | 53.9 | — | |
| Question-Analysis PromptingModel=GPT-3.5 Turbo, Word Count Constraint (n)=1002024.07 | 53.9 | — | |
| CoconutBackbone=LLaMA-3.1-8B-Instruct2025.02 | 53.15 | — | |
| Chain-of-ThoughtModel=GPT-3.5 Turbo, Prompt=CoT2024.07 | 53.1 | — | |
| EoTModel=Llama-3.3-8B2026.03 | 53.01 | — |