Mathematical Reasoning on MultiArith
100AccuracyPrevious SOTA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Previous SOTAPHP=false, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 100 | — | — | — | |
| VanillaBase model=GPT-OSS-120B, Ada.=false, DIM.=false, Dec.=false2026.01 | 100 | — | — | — | |
| LLM-DebateBase model=GPT-OSS-120B, Ada.=false, DIM.=false, Dec.=false2026.01 | 100 | — | — | — | |
| GPTSwarmBase model=GPT-OSS-120B, Ada.=true, DIM.=false, Dec.=false2026.01 | 100 | — | — | — | |
| AgentDropoutBase model=GPT-OSS-120B, Ada.=true, DIM.=false, Dec.=true2026.01 | 100 | — | — | — | |
| G-DesignerBase model=GPT-OSS-120B, Ada.=true, DIM.=true, Dec.=false2026.01 | 100 | — | — | — | |
| TopoDIMBase model=GPT-OSS-120B, Ada.=true, DIM.=true, Dec.=true2026.01 | 100 | — | — | — | |
| VanillaBase model=DeepSeek-V3.2-251201:671B, Ada.=false, DIM.=false, Dec.=false2026.01 | 100 | — | — | — | |
| CoTBase model=DeepSeek-V3.2-251201:671B, Ada.=false, DIM.=false, Dec.=false2026.01 | 100 | — | — | — | |
| LLM-DebateBase model=DeepSeek-V3.2-251201:671B, Ada.=false, DIM.=false, Dec.=false2026.01 | 100 | — | — | — | |
| GPTSwarmBase model=DeepSeek-V3.2-251201:671B, Ada.=true, DIM.=false, Dec.=false2026.01 | 100 | — | — | — | |
| AgentDropoutBase model=DeepSeek-V3.2-251201:671B, Ada.=true, DIM.=false, Dec.=true2026.01 | 100 | — | — | — | |
| G-DesignerBase model=DeepSeek-V3.2-251201:671B, Ada.=true, DIM.=true, Dec.=false2026.01 | 100 | — | — | — | |
| TopoDIMBase model=DeepSeek-V3.2-251201:671B, Ada.=true, DIM.=true, Dec.=true2026.01 | 100 | — | — | — | |
| BiRouterDyn.=true, Dis.=true2025.11 | 100 | — | — | — | |
| Self-consistency (Code-davinci-002)Base Model=Code-davinci-0022024.03 | 100 | — | — | — | |
| TokenSkipBackbone=Qwen3-1.7B, Ratio=1.02026.02 | 99.4 | — | 323 | 1 | |
| Extra-CoT (CHRPO)Backbone=Qwen3-1.7B, Ratio=<POLICY>2026.02 | 99.4 | — | 90 | 0.28 | |
| PAL (Codex)Base Model=Codex2024.03 | 99.2 | — | — | — | |
| Few-Shot-CoT + RCIBase Model=GPT-3.5-Turbo2024.03 | 99.2 | — | — | — | |
| OFA-MAS (pre-trained)Training protocol=One-for-All Topology Design (Pre-trained)2026.01 | 99.11 | — | — | — | |
| OFA-MAS (fine-tuned)Training protocol=One-for-All Topology Design (Fine-tuned)2026.01 | 99.11 | — | — | — | |
| TokenSkipBackbone=Qwen3-1.7B, Ratio=0.82026.02 | 98.9 | — | 266 | 0.83 | |
| Extra-CoTBackbone=Qwen3-1.7B, Ratio=0.82026.02 | 98.9 | — | 278 | 0.87 | |
| TopoDIMBase model=Gemma-3-it:12B, Ada.=true, DIM.=true, Dec.=true2026.01 | 98.85 | — | — | — | |
| MaASDyn.=true, Dis.=false2025.11 | 98.8 | — | — | — | |
| Few-shot-CoT + SCBase Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot-CoT, Self-Consistency=true2024.03 | 98.7 | — | — | — | |
| Few-shot-CoT-CP (GPT-4)Base Model=GPT-4, Prompting Strategy=Few-shot-CoT-CP2024.03 | 98.7 | — | — | — | |
| G-DesignerBase model=Gemma-3-it:12B, Ada.=true, DIM.=true, Dec.=false2026.01 | 98.45 | — | — | — | |
| G-DesignerDyn.=false, Dis.=false2025.11 | 98.33 | — | — | — | |
| Base ModelBackbone=Qwen3-1.7B2026.02 | 98.3 | — | 320 | — | |
| Extra-CoTBackbone=Qwen3-1.7B, Ratio=1.02026.02 | 98.3 | — | 328 | 1 | |
| Extra-CoTBackbone=Qwen3-1.7B, Ratio=0.62026.02 | 98.3 | — | 214 | 0.67 | |
| Zero-shot-CP + SCBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CP, Self-Consistency=true2024.03 | 98.3 | — | — | — | |
| Few-shot-CoT (GPT-4)Base Model=GPT-4, Prompting Strategy=Few-shot-CoT2024.03 | 98.3 | — | — | — | |
| AgentDropoutBase model=Gemma-3-it:12B, Ada.=true, DIM.=false, Dec.=true2026.01 | 98.2 | — | — | — | |
| GPT-4 + PHPPHP=true, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 98.1 | 2.0033 | — | — | |
| GPT-3.5 Turbo + PHPPHP=true, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 98 | 2.0133 | — | — | |
| Few-shot-CoTBase Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot-CoT2024.03 | 98 | — | — | — | |
| GPTSwarmBase model=Gemma-3-it:12B, Ada.=true, DIM.=false, Dec.=false2026.01 | 97.85 | — | — | — | |
| GPT-4PHP=false, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 97.8 | — | — | — | |
| Extra-CoTBackbone=Qwen3-1.7B, Ratio=0.22026.02 | 97.8 | — | 116 | 0.36 | |
| Extra-CoTBackbone=Qwen3-1.7B, Ratio=<POLICY>2026.02 | 97.8 | — | 168 | 0.53 | |
| LLM-DebateBase model=Gemma-3-it:12B, Ada.=false, DIM.=false, Dec.=false2026.01 | 97.62 | — | — | — | |
| GPT-3.5 TurboPHP=false, Prompting Strategy=Complex CoT, Decoding Strategy=greedy2023.04 | 97.5 | — | — | — | |
| AgentVerseDyn.=true, Dis.=false2025.11 | 97.5 | — | — | — | |
| Few-shot-CPBase Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot-CP2024.03 | 97.5 | — | — | — | |
| Few-shot-CoT-CP (GPT-4) + SCBase Model=GPT-4, Prompting Strategy=Few-shot-CoT-CP, Self-Consistency=true2024.03 | 97.5 | — | — | — | |
| LLM-DebateDyn.=false, Dis.=false2025.11 | 97.33 | — | — | — | |
| LLM-BlenderDyn.=false, Dis.=false2025.11 | 97.29 | — | — | — | |
| Zero-Shot-CoT + RCIBase Model=GPT-3.5-Turbo2024.03 | 97.2 | — | — | — | |
| DyLANDyn.=true, Dis.=false2025.11 | 97.12 | — | — | — | |
| CoTBase model=Gemma-3-it:12B, Ada.=false, DIM.=false, Dec.=false2026.01 | 97.01 | — | — | — | |
| Single-agent2025.11 | 96.85 | — | — | — | |
| EIB-LEARNERTraining protocol=Dataset-Specific Training2026.01 | 96.83 | — | — | — | |
| Zero-shot-CoT + SCBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CoT, Self-Consistency=true2024.03 | 96.8 | — | — | — | |
| GPTSwarmDyn.=false, Dis.=false2025.11 | 96.79 | — | — | — | |
| Extra-CoTBackbone=Qwen3-1.7B, Ratio=0.42026.02 | 96.7 | — | 173 | 0.54 | |
| ComplexCoT2025.11 | 96.7 | — | — | — | |
| VanillaBase model=Gemma-3-it:12B, Ada.=false, DIM.=false, Dec.=false2026.01 | 96.68 | — | — | — | |
| SC (CoTx5)2025.11 | 96.58 | — | — | — | |
| G-designerTraining protocol=Dataset-Specific Training2026.01 | 96.5 | — | — | — | |
| LLM-DebateAgent topology=Fixed Multi-Agent (LLM-Debate)2026.01 | 96.36 | — | — | — | |
| CoT2025.11 | 96.31 | — | — | — | |
| Zero-shot-CoT-CPBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CoT-CP2024.03 | 96.2 | — | — | — | |
| TokenSkipBackbone=Qwen3-1.7B, Ratio=0.62026.02 | 96.1 | — | 260 | 0.81 | |
| MacNetDyn.=false, Dis.=false2025.11 | 96.03 | — | — | — | |
| AgentDropoutTraining protocol=Dataset-Specific Training2026.01 | 95.6 | — | — | — | |
| Zero-shot-CPBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CP2024.03 | 95.2 | — | — | — | |
| Zero-shot-CPPrompting Strategy=Zero-shot-CP, Base Model=GPT-3.5-Turbo, Number of runs=52024.03 | 95.13 | — | — | — | |
| TokenSkipBackbone=Qwen3-1.7B, Ratio=<POLICY>2026.02 | 95 | — | 175 | 0.55 | |
| Zero-shot-CoTPrompting Strategy=Zero-shot-CoT, Base Model=GPT-3.5-Turbo, Number of runs=52024.03 | 94.87 | — | — | — | |
| Zero-shot-CoTBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CoT2024.03 | 94.8 | — | — | — | |
| AgentPruneTraining protocol=Dataset-Specific Training2026.01 | 94.65 | — | — | — | |
| CompleteAgent topology=Fixed Multi-Agent (Complete)2026.01 | 94.53 | — | — | — | |
| CoT-SFTReasoning Type=CoT, Training=SFT2026.01 | 94.4 | — | 19 | — | |
| ATP-LatentReasoning Type=Latent, Training=RL2026.01 | 94.4 | — | 7.1 | — | |
| TokenSkipBackbone=Qwen3-1.7B, Ratio=0.42026.02 | 94.4 | — | 177 | 0.55 | |
| SC (CoT)Prompting strategy=Single-Agent Prompting2026.01 | 94.12 | — | — | — | |
| RandomAgent topology=Fixed Multi-Agent (Random)2026.01 | 94.08 | — | — | — | |
| ATP-Latent -w/o Stop HeadReasoning Type=Latent, Training=RL, Stop Head=Excluded2026.01 | 93.9 | — | 9 | — | |
| CoLTSeed count=22026.02 | 93.9 | 7.27 | — | — | |
| TreeAgent topology=Fixed Multi-Agent (Tree)2026.01 | 93.68 | — | — | — | |
| ChainAgent topology=Fixed Multi-Agent (Chain)2026.01 | 93.27 | — | — | — | |
| CoTPrompting strategy=Single-Agent Prompting2026.01 | 93.25 | — | — | — | |
| CoT2026.02 | 93.2 | 13.7 | — | — | |
| CoT2026.02 | 93.2 | 13.7 | — | — | |
| VanillaPrompting strategy=Direct Answer2026.01 | 93.09 | — | — | — | |
| LSTRCompression ratio=2, Ablation=w/o Skip2026.02 | 92.8 | 7.37 | — | — | |
| ATP-Latent -w/o VAEReasoning Type=Latent, Training=RL, VAE=Excluded2026.01 | 92.8 | — | 7.2 | — | |
| CoLTSeed count=12026.02 | 92.8 | 5.3 | — | — | |
| Zero-shot-PoT (Codex)Base Model=Codex, Prompting Strategy=Zero-shot-PoT2024.03 | 92.2 | — | — | — | |
| CoLaRCompression ratio=22026.02 | 91.3 | 7.35 | — | — | |
| COLARMultiplier=2x2026.02 | 91.3 | 7.35 | — | — | |
| TokenSkipBackbone=Qwen3-1.7B, Ratio=0.22026.02 | 91.1 | — | 108 | 0.34 | |
| ATP-Latent w Only Coherence as RewardReasoning Type=Latent, Training=RL, Reward Mechanism=Unsupervised Coherence2026.01 | 90.6 | — | 7.1 | — | |
| LSTRCompression ratio=22026.02 | 90.5 | 7.44 | — | — | |
| ATP-Latent -w/o RL (Ours-SFT)Reasoning Type=Latent, Training=SFT2026.01 | 90 | — | 7.1 | — | |
| Zero-shot-CoT + self consistencyModel=PaLM (540B), Prompting Strategy=Zero-shot, Chain-of-thought=true, Self-consistency=40 paths2022.05 | 89 | — | — | — | |
| LSTRCompression ratio=52026.02 | 88.2 | 3.17 | — | — |