Mathematical Reasoning on GSM8K (HS, Acc)
96.81Accuracy (Acc)SIGMA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SIGMAMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=full, Backbone=DeepSeek-V3.22026.05 | 96.81 | — | 93.91 | — | — | |
| G-DesignerMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 96.66 | — | 93.22 | — | — | |
| CPMajModel=Average across models2026.04 | 96.1 | — | — | — | — | |
| Qwen3Parameters=8B, Architecture=AR2026.04 | 96 | — | — | — | — | |
| CPMaxModel=Average across models2026.04 | 96 | — | — | — | — | |
| CPAgrModel=Average across models2026.04 | 96 | — | — | — | — | |
| SCCoTModel=Average across models2026.04 | 95.7 | — | — | — | — | |
| GoAMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=partial, Backbone=DeepSeek-V3.22026.05 | 95.43 | — | 92.86 | — | — | |
| GPTSwarmMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 95.24 | — | 92.39 | — | — | |
| ComplexCoTMulti-agent collaboration (Mul.)=no, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 95.03 | — | 91.01 | — | — | |
| I-DLMParameters=8B2026.04 | 95 | — | — | — | — | |
| CoTMulti-agent collaboration (Mul.)=no, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 94.85 | — | 90.68 | — | — | |
| Athena-PRMPolicy Model=Qwen2.5-7B, Selection Strategy=Athena-PRM, N=82025.06 | 94.8 | — | — | — | — | |
| MoAMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 94.77 | — | 89.83 | — | — | |
| VisualPRM-8BPolicy Model=Qwen2.5-7B, Selection Strategy=VisualPRM-8B, N=82025.06 | 94.5 | — | — | — | — | |
| VanillaMulti-agent collaboration (Mul.)=no, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 94.47 | — | 87.69 | — | — | |
| Athena-ORMPolicy Model=Qwen2.5-7B, Selection Strategy=Athena-ORM, N=82025.06 | 94 | — | — | — | — | |
| SCPoTModel=Average across models2026.04 | 93.9 | — | — | — | — | |
| D2EvoBackbone=Qwen3-8B-Base, #Data=0.4K, Iteration=Iter 32026.05 | 93.7 | — | — | — | — | |
| D2EvoBackbone=Qwen3-8B-Base, #Data=0.3K, Iteration=Iter 22026.05 | 93.68 | — | — | — | — | |
| Full DataBackbone=Qwen3-8B-Base, #Data=19K2026.05 | 93.56 | — | — | — | — | |
| R-ZeroBackbone=Qwen3-8B-Base, Iteration=Iter 22026.05 | 93.56 | — | — | — | — | |
| R-ZeroBackbone=Qwen3-8B-Base, Iteration=Iter 32026.05 | 93.51 | — | — | — | — | |
| R-ZeroBackbone=Qwen3-8B-Base, Iteration=Iter 12026.05 | 93.43 | — | — | — | — | |
| Self-consistencyPolicy Model=Qwen2.5-7B, Selection Strategy=Self-consistency, N=82025.06 | 93.4 | — | — | — | — | |
| D2EvoBackbone=Qwen3-8B-Base, #Data=1K, Iteration=Iter 12026.05 | 93.38 | — | — | — | — | |
| CoTModel=Average across models2026.04 | 93.2 | — | — | — | — | |
| Qwen3-4B-Base + Resample w/ LOPE (w/o Training Signal Shaping)Model=Qwen3-4B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=false2026.05 | 92.95 | — | — | — | — | |
| Qwen3-4B-Base + Resample w/ LOPE (w/ Training Signal Shaping)Model=Qwen3-4B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=true2026.05 | 92.95 | — | — | — | — | |
| Qwen3-4B-Base + Resample w/ Naive PromptModel=Qwen3-4B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=Naive Prompt2026.05 | 92.87 | — | — | — | — | |
| SPICEBackbone=Qwen3-4B-Base, #Data=20K2026.05 | 92.7 | — | — | — | — | |
| SPICEBackbone=Qwen3-8B-Base, #Data=20K2026.05 | 92.7 | — | — | — | — | |
| D2EvoBackbone=Qwen3-4B-Base, #Data=0.3K, Iteration=Iter 22026.05 | 92.5 | — | — | — | — | |
| D2EvoBackbone=Qwen3-4B-Base, #Data=0.1K, Iteration=Iter 32026.05 | 92.46 | — | — | — | — | |
| D2EvoBackbone=Qwen3-4B-Base, #Data=1K, Iteration=Iter 12026.05 | 92.4 | — | — | — | — | |
| AZRBackbone=Qwen3-8B-Base2026.05 | 92.2 | — | — | — | — | |
| Base ModelBackbone=Qwen3-8B-Base2026.05 | 92.19 | — | — | — | — | |
| SupervisedLLM=qwen2.5-7b-instruct2026.05 | 92.11 | — | — | — | — | |
| R-ZeroBackbone=Qwen3-4B-Base, Iteration=Iter 22026.05 | 92.11 | — | — | — | — | |
| R-ZeroBackbone=Qwen3-4B-Base, Iteration=Iter 32026.05 | 92.11 | — | — | — | — | |
| R-ZeroBackbone=Qwen3-4B-Base, Iteration=Iter 12026.05 | 91.88 | — | — | — | — | |
| PoTModel=Average across models2026.04 | 91.8 | — | — | — | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 0, Retention ratio (r)=0.92026.03 | 91.74 | — | — | — | — | |
| Qwen3-4B-Base + GRPOModel=Qwen3-4B-Base, Optimization Method=GRPO2026.05 | 91.74 | — | — | — | — | |
| Qwen2.5-7BPolicy Model=Qwen2.5-7B, Selection Strategy=Zero-shot, N=12025.06 | 91.6 | — | — | — | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 0, Retention ratio (r)=0.752026.03 | 91.51 | — | — | — | — | |
| LOVER (sum)LLM=qwen2.5-7b-instruct2026.05 | 91.43 | — | — | — | — | |
| H-GRPOBackbone=Qwen2.5-7B2026.05 | 91.43 | — | — | — | — | |
| Jacobi ForcingParameters=7B2026.04 | 91.4 | — | — | — | — | |
| Full DataBackbone=Qwen3-4B-Base, #Data=19K2026.05 | 91.36 | — | — | — | — | |
| Long2ShortBackbone=Qwen2.5-7B2026.05 | 91.13 | — | — | — | — | |
| NBDiffParameters=7B2026.04 | 91 | — | — | — | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 2, Retention ratio (r)=0.752026.03 | 90.9 | — | — | — | — | |
| BaseBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 90.64 | — | — | — | — | |
| BaseModel=Qwen-2.5-7B-Inst, Training Strategy=Base, Evaluation Protocol=0-shot2026.04 | 90.64 | — | — | — | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 2, Retention ratio (r)=0.92026.03 | 90.52 | — | — | — | — | |
| LightningRLParameters=8B2026.04 | 90.3 | — | — | — | — | |
| Qwen2.5-Math-7B + Resample w/ LOPE (w/ Training Signal Shaping)Model=Qwen2.5-Math-7B, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=true2026.05 | 90.3 | — | — | — | — | |
| POPBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 90.24 | — | — | — | — | |
| WeDLMParameters=8B2026.04 | 90.2 | — | — | — | — | |
| POPModel=Qwen-2.5-7B-Inst, Training Strategy=POP, Evaluation Protocol=0-shot2026.04 | 90.18 | — | — | — | — | |
| CoT-Decoding (sum)LLM=qwen2.5-7b-instruct2026.05 | 89.76 | — | — | — | — | |
| GRPOBackbone=Qwen2.5-7B2026.05 | 89.61 | — | — | — | — | |
| Majority VotingLLM=qwen2.5-7b-instruct2026.05 | 89.46 | — | — | — | — | |
| BaseModel=Qwen2.5-7B-Instruct2025.08 | 89.4 | — | — | — | — | |
| CosFnBackbone=Qwen2.5-7B2026.05 | 89.31 | — | — | — | — | |
| AZRBackbone=Qwen3-4B-Base2026.05 | 89.3 | — | — | — | — | |
| TraceLiftModel=Qwen3-4B2026.05 | 89.23 | — | — | — | — | |
| Qwen3-4BVersion=25072026.04 | 89.16 | — | — | — | — | |
| TraceLiftModel=Qwen2.5-7B2026.05 | 89.16 | — | — | — | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 0, Retention ratio (r)=1.02026.03 | 89.08 | — | — | — | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 2, Retention ratio (r)=1.02026.03 | 89.08 | — | — | — | — | |
| Exec-onlyModel=Qwen3-4B2026.05 | 89.01 | — | — | — | — | |
| Train on DBackbone=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 88.87 | — | — | — | — | |
| BaseModel=Qwen3-4B2026.05 | 88.48 | — | — | — | — | |
| Base ModelBackbone=Qwen3-4B-Base2026.05 | 88.15 | — | — | — | — | |
| Train on DModel=Qwen-2.5-7B-Inst, Training Strategy=Naive pretraining on D, Evaluation Protocol=0-shot2026.04 | 87.98 | — | — | — | — | |
| TokenBuncherModel=Qwen2.5-7B-Instruct2025.08 | 87.6 | — | — | — | — | |
| Exec-onlyModel=Qwen2.5-7B2026.05 | 87.04 | — | — | — | — | |
| Qwen3-8BType=AR2026.06 | 86.5 | — | — | — | — | |
| D2EvoBackbone=Llama-3.1-8B-Instruct, #Data=0.6K, Iteration=Iter 12026.05 | 86.4 | — | — | — | — | |
| D2EvoBackbone=Llama-3.1-8B-Instruct, #Data=0.2K, Iteration=Iter 22026.05 | 86.4 | — | — | — | — | |
| Qwen2.5-Math-7B + Resample w/ LOPE (w/o Training Signal Shaping)Model=Qwen2.5-Math-7B, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=false2026.05 | 86.35 | — | — | — | — | |
| D2EvoBackbone=Llama-3.1-8B-Instruct, #Data=0.4K, Iteration=Iter 32026.05 | 86.3 | — | — | — | — | |
| SPGSeq Len=256, Shot count=0-SHOT2026.04 | 86.1 | — | — | — | — | |
| Full DataBackbone=Llama-3.1-8B-Instruct, #Data=19K2026.05 | 85.8 | — | — | — | — | |
| Qwen2.5-Math-7B + GRPOModel=Qwen2.5-Math-7B, Optimization Method=GRPO2026.05 | 85.06 | — | — | — | — | |
| Gemma-3-4B2026.04 | 84.76 | — | — | — | — | |
| SPGSeq Len=512, Shot count=0-SHOT2026.04 | 84.5 | — | — | — | — | |
| OPTIMERBase Model=Gemma 3 27B2026.03 | 84.38 | — | — | — | — | |
| R-ZeroBackbone=Llama-3.1-8B-Instruct2026.05 | 84.23 | — | — | — | — | |
| AZRBackbone=Llama-3.1-8B-Instruct2026.05 | 83.98 | — | — | — | — | |
| Base ModelBackbone=Llama-3.1-8B-Instruct2026.05 | 83.7 | — | — | — | — | |
| Qwen3-1.7B-Base + Resample w/ LOPE (w/o Training Signal Shaping)Model=Qwen3-1.7B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=false2026.05 | 83.55 | — | — | — | — | |
| LOVER (max)LLM=qwen2.5-7b-instruct2026.05 | 83.47 | — | — | — | — | |
| DreamReasoner-8BType=Block Diffusion2026.06 | 83.4 | — | — | — | — | |
| POPBackbone=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | 83.26 | — | — | — | — | |
| SupervisedLLM=llama-3.1-8b-instruct2026.05 | 83.24 | — | — | — | — | |
| DTMSeq Len=512, Shot count=0-SHOT2026.04 | 83.2 | — | — | — | — | |
| H-GRPOBackbone=Qwen2.5-3B2026.05 | 83.09 | — | — | — | — |