Mathematical Reasoning on AQuA (Accuracy and Improvement)
87.01AccuracyADv2
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ADv2Backbone=Qwen3-14B, Indicator Pool Source=Qwen3-8B-based MAS2026.02 | 87.01 | — | |
| Dynamic-MASBackbone=Qwen3-14B2026.02 | 86.22 | — | |
| Dynamic-MAS + Multi-TAGFramework=Dynamic-MAS, Technique=Multi-TAG2026.02 | 85.83 | — | |
| Dynamic-MASFramework=Dynamic-MAS2026.02 | 85.43 | — | |
| Dynamic-MAS + PRMFramework=Dynamic-MAS, Technique=PRM2026.02 | 85.43 | — | |
| Fixed-MASFramework=Fixed-MAS2026.02 | 85.04 | — | |
| Dynamic-MAS + Self-RefineFramework=Dynamic-MAS, Technique=Self-Refine2026.02 | 85.04 | — | |
| Single Agent + CoTFramework=Single Agent, Technique=Chain-of-Thought2026.02 | 84.65 | — | |
| Fixed-MAS + ADv2Framework=Fixed-MAS, Technique=ADv22026.02 | 84.65 | — | |
| Fixed-MAS + PRMFramework=Fixed-MAS, Technique=PRM2026.02 | 84.25 | — | |
| Single AgentFramework=Single Agent2026.02 | 83.86 | — | |
| Fixed-MAS + ADv1Framework=Fixed-MAS, Technique=ADv12026.02 | 83.86 | — | |
| Dynamic-MAS + ADv2Framework=Dynamic-MAS, Technique=ADv22026.02 | 83.86 | — | |
| Single AgentBackbone=Qwen3-14B2026.02 | 83.86 | — | |
| Fixed-MAS + Self-RefineFramework=Fixed-MAS, Technique=Self-Refine2026.02 | 83.46 | — | |
| Fixed-MAS + Multi-TAGFramework=Fixed-MAS, Technique=Multi-TAG2026.02 | 83.07 | — | |
| MOCEdge Density=0.3, Backbone=Gemma-2-27B-Instruct2026.06 | 74.41 | — | |
| MOCEdge Density=0.5, Backbone=Gemma-2-27B-Instruct2026.06 | 73.23 | — | |
| MOCEdge Density=1.0, Backbone=Gemma-2-27B-Instruct2026.06 | 72.83 | — | |
| Vanilla MASEdge Density=0.5, Backbone=Gemma-2-27B-Instruct2026.06 | 72.44 | — | |
| MOCEdge Density=0.7, Backbone=Gemma-2-27B-Instruct2026.06 | 72.05 | — | |
| Vanilla MASEdge Density=0.7, Backbone=Gemma-2-27B-Instruct2026.06 | 70.87 | — | |
| Vanilla MASEdge Density=1.0, Backbone=Gemma-2-27B-Instruct2026.06 | 70.87 | — | |
| Vanilla MASEdge Density=0.3, Backbone=Gemma-2-27B-Instruct2026.06 | 69.69 | — | |
| Single LLMBackbone=Gemma-2-27B-Instruct2026.06 | 65.35 | — | |
| Complex CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Complex CoT, Enhancement=RICP2024.07 | 41.12 | 6.78 | |
| Few-shot CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=RICP2024.07 | 40.19 | 4.21 | |
| Auto CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Auto CoT, Enhancement=RICP2024.07 | 39.02 | 2.57 | |
| Few-shot CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=RICP2024.07 | 38.55 | 3.74 | |
| Zero-shot CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=RICP2024.07 | 38.08 | 4.44 | |
| Zero-shot CoT with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=RICP2024.07 | 38.08 | 7.01 | |
| Complex CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Complex CoT, Enhancement=RICP2024.07 | 37.85 | 5.61 | |
| Auto CoT with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Auto CoT, Enhancement=RICP2024.07 | 37.15 | 0.7 | |
| PartitionSelBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 36.61 | — | |
| Auto CoTLLM Model=Qwen-Turbo, Prompting Strategy=Auto CoT, Enhancement=Vanilla2024.07 | 36.45 | — | |
| Auto CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Auto CoT, Enhancement=Vanilla2024.07 | 36.45 | — | |
| Few-shot CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=Vanilla2024.07 | 35.98 | — | |
| Few-shot CoTLLM Model=Qwen-Turbo, Prompting Strategy=Few-shot CoT, Enhancement=Vanilla2024.07 | 34.81 | — | |
| FRAMEBackbone=LLaMA-3.1-8B2026.06 | 34.8 | — | |
| Complex CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Complex CoT, Enhancement=Vanilla2024.07 | 34.35 | — | |
| IDBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 34.25 | — | |
| MoABackbone=LLaMA-3.1-8B2026.06 | 34 | — | |
| FourierMoEBackbone=LLaMA-3.1-8B2026.06 | 33.8 | — | |
| Zero-shot CoTLLM Model=Qwen-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=Vanilla2024.07 | 33.64 | — | |
| FlyLoRABackbone=LLaMA-3.1-8B2026.06 | 33.6 | — | |
| HMoRABackbone=LLaMA-3.1-8B2026.06 | 33.3 | — | |
| MixLoRABackbone=LLaMA-3.1-8B2026.06 | 32.9 | — | |
| HydraLoRABackbone=LLaMA-3.1-8B2026.06 | 32.4 | — | |
| Complex CoTLLM Model=Qwen-Turbo, Prompting Strategy=Complex CoT, Enhancement=Vanilla2024.07 | 32.24 | — | |
| DoRABackbone=LLaMA-3.1-8B2026.06 | 31.8 | — | |
| COLMBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 31.5 | — | |
| GradNormBackbone=Llama-3.1-8B, Training Dataset=MetaMathQA, Subset Ratio=12.5%2026.06 | 31.5 | — | |
| FourierFTBackbone=LLaMA-3.1-8B2026.06 | 31.4 | — | |
| LoRABackbone=LLaMA-3.1-8B2026.06 | 31.2 | — | |
| Zero-shot CoTLLM Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot CoT, Enhancement=Vanilla2024.07 | 31.07 | — | |
| Standard Prompting with RICPLLM Model=GPT-3.5-Turbo, Prompting Strategy=Standard Prompting, Enhancement=RICP2024.07 | 25 | 4.67 | |
| Standard PromptingLLM Model=GPT-3.5-Turbo, Prompting Strategy=Standard Prompting, Enhancement=Vanilla2024.07 | 20.33 | — | |
| Standard Prompting with RICPLLM Model=Qwen-Turbo, Prompting Strategy=Standard Prompting, Enhancement=RICP2024.07 | 17.76 | 1.87 | |
| Standard PromptingLLM Model=Qwen-Turbo, Prompting Strategy=Standard Prompting, Enhancement=Vanilla2024.07 | 15.89 | — |