Math Reasoning on MATH (test) (Accuracy)
96AccuracyMAD
Evaluation Results
| Method | Links | |
|---|---|---|
| MADCategory=Multi-Agent2026.05 | 96 | |
| MoACategory=Multi-Agent2026.05 | 95.6 | |
| DMADCategory=Multi-Agent2026.05 | 95 | |
| AgentPSOCategory=Multi-Agent2026.05 | 94.6 | |
| Self-RefineCategory=Advanced Single-Agent2026.05 | 91.6 | |
| ReflectionCategory=Advanced Single-Agent2026.05 | 90.8 | |
| Chain-of-ThoughtCategory=Vanilla Single-Agent2026.05 | 90.4 | |
| Step-Back PromptingCategory=Vanilla Single-Agent2026.05 | 90.4 | |
| Self-ConsistencyCategory=Advanced Single-Agent2026.05 | 88.8 | |
| AoTBackbone=GPT-4o-mini2026.04 | 83.6 | |
| Agent-GWOBackbone=Gemma-3-12b-it2026.04 | 82.1 | |
| ToTCategory=Advanced Single-Agent2026.05 | 81.6 | |
| GoTBackbone=Gemma-3-12b-it2026.04 | 80.3 | |
| Agent-GWOBackbone=GPT-4o-mini2026.04 | 80.2 | |
| AoTBackbone=Gemma-3-12b-it2026.04 | 79.6 | |
| ToTBackbone=Gemma-3-12b-it2026.04 | 79.5 | |
| AFlowBackbone=Gemma-3-12b-it2026.04 | 79 | |
| AFlowBackbone=GPT-4o-mini2026.04 | 78.9 | |
| GoTBackbone=GPT-4o-mini2026.04 | 78.6 | |
| ToTBackbone=GPT-4o-mini2026.04 | 77.8 | |
| Self-RefineBackbone=Gemma-3-12b-it2026.04 | 77.6 | |
| CoT-SC/n=5Backbone=GPT-4o-mini2026.04 | 76.3 | |
| Self-RefineBackbone=GPT-4o-mini2026.04 | 75.1 | |
| CoTBackbone=GPT-4o-mini2026.04 | 74.8 | |
| CoT-SC/n=5Backbone=Gemma-3-12b-it2026.04 | 74.6 | |
| Agent-GWOBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 74.1 | |
| CoTBackbone=Gemma-3-12b-it2026.04 | 72.8 | |
| AoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 72 | |
| AFlowBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 71.9 | |
| GoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 71.7 | |
| ToTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 71.3 | |
| Self-RefineBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 70.4 | |
| CoT-SC/n=5Backbone=Qwen2.5-Coder-7B-Instruct2026.04 | 67.3 | |
| CoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 65.8 | |
| MENTORCOLLAB MLPGenerator=Qwen3-8B-Base, Mentor=Qwen3-32B, rho=25%, Decoding Strategy=greedy, Token Budget=5122026.02 | 46.8 | |
| OMACOptimization Dimension Category=Functional Dimension, Specific Dimension Optimized=Fun-1.12025.05 | 35.17 | |
| OMACOptimization Dimension Category=Functional Dimension, Specific Dimension Optimized=Fun-1.22025.05 | 34.91 | |
| OMACOptimization Dimension Category=Functional Dimension, Specific Dimension Optimized=Fun-22025.05 | 33.95 | |
| OMACOptimization Dimension Category=Structural Dimension, Specific Dimension Optimized=Str-32025.05 | 33.7 | |
| OMACOptimization Dimension Category=Structural Dimension, Specific Dimension Optimized=Str-22025.05 | 33.41 | |
| OMACOptimization Dimension Category=Structural Dimension, Specific Dimension Optimized=Str-12025.05 | 33.34 | |
| AFlowMethod Category=Baselines2025.05 | 32.49 | |
| DyLANMethod Category=Baselines2025.05 | 32.35 | |
| LLM DebateMethod Category=Baselines2025.05 | 29.42 | |
| ADASMethod Category=Baselines2025.05 | 28.94 | |
| SEMethod Category=Baselines2025.05 | 28.72 | |
| Co-LLMGenerator=Qwen3-8B-Base, Mentor=Qwen3-14B2026.02 | 25 | |
| MENTORCOLLAB MLPGenerator=Gemma-3-4B-PT, Mentor=Qwen3-14B, rho=25%, Decoding Strategy=greedy, Token Budget=5122026.02 | 21 | |
| MENTORCOLLAB MLPGenerator=Llama3.1-8B, Mentor=Qwen3-32B, rho=25%, Decoding Strategy=greedy, Token Budget=5122026.02 | 18 | |
| MENTORCOLLAB FREEGenerator=Gemma-3-4B-PT, Mentor=Qwen3-14B, rho=25%, Decoding Strategy=greedy, Token Budget=5122026.02 | 15.8 | |
| Generator BaselineModel=Gemma-3-4B-PT, Decoding Strategy=greedy, Token Budget=5122026.02 | 14.2 | |
| NudgingGenerator=Qwen3-1.7B, Mentor=R1-Distilled-Llama-70B, gamma=0.4, Decoding Strategy=greedy, Token Budget=5122026.02 | 13.6 | |
| FedMomentumBackbone=LLaMA2-7B, LoRA rank=32, scaling factor=64, communication rounds=502026.03 | 4.44 | |
| RoLoRABackbone=LLaMA2-7B, LoRA rank=32, scaling factor=64, communication rounds=502026.03 | 4.42 | |
| FFA-LoRABackbone=LLaMA2-7B, LoRA rank=32, scaling factor=64, communication rounds=502026.03 | 4.18 | |
| FedEx-LoRABackbone=LLaMA2-7B, LoRA rank=32, scaling factor=64, communication rounds=502026.03 | 3.97 | |
| FLoRABackbone=LLaMA2-7B, LoRA rank=32, scaling factor=64, communication rounds=502026.03 | 3.86 | |
| FedITBackbone=LLaMA2-7B, LoRA rank=32, scaling factor=64, communication rounds=502026.03 | 3.84 | |
| Pre-trainedBackbone=LLaMA2-7B2026.03 | 2.5 |