Mathematical Reasoning on GSM8K (Accuracy & Accuracy (POT))
89.15AccuracyGroup-Debate
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Group-DebateAggregation=AgentAuditor2026.02 | 89.15 | — | |
| GPTSwarmAggregation=AgentAuditor2026.02 | 88.15 | — | |
| Qwen3-4BParams=4B2025.12 | 87.79 | — | |
| DyLanAggregation=AgentAuditor2026.02 | 87.58 | — | |
| LLM-DebateAggregation=AgentAuditor2026.02 | 87.43 | — | |
| AgentPruneAggregation=AgentAuditor2026.02 | 87.34 | — | |
| GPTSwarmAggregation=LLM-as-Judge2026.02 | 86.05 | — | |
| AgentPruneAggregation=LLM-as-Judge2026.02 | 85.83 | — | |
| LLM-DebateAggregation=LLM-as-Judge2026.02 | 85.37 | — | |
| Group-DebateAggregation=LLM-as-Judge2026.02 | 85.36 | — | |
| GPTSwarmAggregation=MV2026.02 | 84.89 | — | |
| DyLanAggregation=LLM-as-Judge2026.02 | 84.63 | — | |
| AgentPruneAggregation=MV2026.02 | 84.38 | — | |
| Group-DebateAggregation=MV2026.02 | 83.98 | — | |
| LLM-DebateAggregation=MV2026.02 | 83.52 | — | |
| DyLanAggregation=MV2026.02 | 82.03 | — | |
| SC (Self-Consistency)Category=Single-Agent2026.02 | 80.79 | — | |
| Qwen2.5-3BParams=3B2025.12 | 79.1 | — | |
| llama-3.2-3BParams=3B2025.12 | 77.7 | — | |
| Qwen3-1.7BParams=1.7B2025.12 | 75.44 | — | |
| CoTCategory=Single-Agent2026.02 | 74.22 | — | |
| VanillaCategory=Single-Agent2026.02 | 72.76 | — | |
| SMoAModel=Llama-3-8B, Trainable=0.8256, r=32, n=22026.01 | 72.14 | — | |
| HiRAModel=Llama-3-8B, Trainable=0.8256, r=322026.01 | 70.81 | — | |
| MeLoRAModel=Llama-3-8B, Trainable=0.8256, r=32, n=22026.01 | 69.36 | — | |
| Qwen2.5-1.5BParams=1.5B2025.12 | 68.5 | — | |
| SMoAModel=Llama-3-8B, Trainable=0.4128, r=16, n=22026.01 | 68.37 | — | |
| OLMo-2-0425-1BParams=1B2025.12 | 68.3 | — | |
| MoRAModel=Llama-3-8B, Trainable=0.8241, r=322026.01 | 67.89 | — | |
| SmolLM3-3BParams=3B2025.12 | 67.63 | — | |
| HiRAModel=Llama-3-8B, Trainable=0.4128, r=162026.01 | 67.63 | — | |
| MeLoRAModel=Llama-3-8B, Trainable=0.4128, r=16, n=22026.01 | 66.76 | — | |
| YuLan-Mini-2.4BParams=2.4B2025.12 | 66.65 | — | |
| DoRAModel=Llama-3-8B, Trainable=0.8256, r=322026.01 | 66.12 | — | |
| LoRAModel=Llama-3-8B, Trainable=0.8256, r=322026.01 | 65.89 | — | |
| SSMLoRAModel=Llama-3-8B, Trainable=0.8024, r=322026.01 | 64.75 | — | |
| Qwen3-0.6BParams=0.6B2025.12 | 59.59 | — | |
| Qwen2-1.5BParams=1.5B2025.12 | 58.5 | — | |
| PCMind-2.1-Kaiyuan-2BParams=2B2025.12 | 51.33 | — | |
| llama-3.2-1BParams=1B2025.12 | 44.4 | — | |
| LLAMA PRO INSTRUCTTraining Stage=SFT comparison2024.01 | 43.59 | 55.61 | |
| SmolLM2-1.7BParams=1.7B2025.12 | 31.1 | — | |
| gemma2-2BParams=2B2025.12 | 23.9 | — | |
| LLAMA PROParameters=8B, Training Stage=Pretrained comparison2024.01 | 17.89 | 25.42 | |
| PromptTuningModel=Llama-3-8B, Trainable=0.00122026.01 | 15.62 | — | |
| LLaMA2Parameters=7B, Training Stage=Pretrained comparison2024.01 | 14.48 | 17.68 | |
| CrystalCoderParameters=7B, Training Stage=Pretrained comparison2024.01 | 10.77 | 24.96 | |
| StarCoderParameters=15B, Training Stage=Pretrained comparison2024.01 | 9.48 | 25.09 | |
| LLaMAParameters=7B, Training Stage=Pretrained comparison2024.01 | 8.04 | 10.46 | |
| CodeLLaMA-InstructParameters=7B, Training Stage=SFT comparison2024.01 | 7.96 | 34.67 | |
| LLaMA2-ChatParameters=7B, Training Stage=SFT comparison2024.01 | 7.35 | 19.73 | |
| CodeLLaMAParameters=7B, Training Stage=Pretrained comparison2024.01 | 5.16 | 25.2 | |
| WizardCoder-PythonParameters=7B, Training Stage=SFT comparison2024.01 | 4.7 | 17.6 | |
| FalconParameters=7B, Training Stage=Pretrained comparison2024.01 | 4.62 | 4.32 | |
| OpenLLaMA-v2Parameters=7B, Training Stage=Pretrained comparison2024.01 | 3.49 | 5.46 | |
| WizardMathParameters=7B, Training Stage=SFT comparison2024.01 | 2.73 | 25.57 | |
| P-TuningModel=Llama-3-8B, Trainable=0.74282026.01 | 2.65 | — |