Mathematical Reasoning on GSM8K (Accuracy %)
97.8Accuracy (GSM8K)Pass@8 (Upper Bound)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Pass@8 (Upper Bound)Decoding Strategy=Pass@82025.05 | 97.8 | — | — | — | |
| RLHFlow-PRM-Deepseek-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 96.7 | — | — | — | |
| SCOPEDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 96.7 | — | — | — | |
| SCOPE (Step Skipping)Decoding Strategy=Best-of-8 strategy, Ablation=Step Skipping2025.05 | 96.7 | — | — | — | |
| SCOPE (w/o code translation)Decoding Strategy=Best-of-8 strategy, Ablation=w/o code translation2025.05 | 96.6 | — | — | — | |
| SCOPE (Step Replacement)Decoding Strategy=Best-of-8 strategy, Ablation=Step Replacement2025.05 | 96.6 | — | — | — | |
| Majority @8Decoding Strategy=Majority @82025.05 | 96.5 | — | — | — | |
| Qwen2.5-Math-7B-PRM800KDecoding Strategy=Best-of-8 strategy, Training Data=manually annotated (PRM800K)2025.05 | 96.5 | — | — | — | |
| SCOPE (w/o AST normalization)Decoding Strategy=Best-of-8 strategy, Ablation=w/o AST normalization2025.05 | 96.5 | — | — | — | |
| Skywork-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 96.4 | — | — | — | |
| NVFP4Backbone=LongCat 560B2026.02 | 96.29 | — | — | 0.38 | |
| Math-Shepherd-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 96.2 | — | — | — | |
| Skywork-PRM-1.5BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 96.2 | — | — | — | |
| RLHFlow-PRM-Mistral-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 96 | — | — | — | |
| BF16Backbone=LongCat 560B2026.02 | 95.91 | — | — | — | |
| EurusPRM-Stage2Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 95.8 | — | — | — | |
| EFFGENBackbone=Qwen2.5-32B-Instruct2026.01 | 95.75 | — | — | — | |
| HiF4Backbone=DeepSeek-V3.1 671B2026.02 | 95.75 | — | — | 1.29 | |
| HiF4Backbone=LongCat 560B2026.02 | 95.75 | — | — | -0.16 | |
| Raw ModelBackbone=Qwen2.5-32B-Instruct2026.01 | 95.6 | — | — | — | |
| GreedyDecoding Strategy=Greedy2025.05 | 95.5 | — | — | — | |
| NVFP4+PTSBackbone=LongCat 560B2026.02 | 95.45 | — | — | -0.46 | |
| NVFP4+PTSBackbone=DeepSeek-V3.1 671B2026.02 | 95.3 | — | — | 0.84 | |
| EurusPRM-Stage1Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 95.1 | — | — | — | |
| LangChainBackbone=Qwen2.5-32B-Instruct2026.01 | 95.07 | — | — | — | |
| NVFP4Backbone=DeepSeek-V3.1 671B2026.02 | 95 | — | — | 0.54 | |
| Self-Consistency2026.02 | 95 | — | — | — | |
| EFFGENBackbone=Qwen2.5-14B-Instruct2026.01 | 94.84 | — | — | — | |
| AutogenBackbone=Qwen2.5-32B-Instruct2026.01 | 94.54 | — | — | — | |
| BF16Backbone=DeepSeek-V3.1 671B2026.02 | 94.46 | — | — | — | |
| AutogenBackbone=Qwen2.5-14B-Instruct2026.01 | 94.09 | — | — | — | |
| Primitives-based MAS2026.02 | 93.8 | — | — | — | |
| Raw ModelBackbone=Qwen2.5-14B-Instruct2026.01 | 93.78 | — | — | — | |
| SmolagentsBackbone=Qwen2.5-32B-Instruct2026.01 | 93.4 | — | — | — | |
| AgentVerse2026.02 | 93.4 | — | — | — | |
| MAS-GPT2026.02 | 93.4 | — | — | — | |
| GPTSwarm2026.02 | 93.2 | — | — | — | |
| Quality-Diversity2026.02 | 93 | — | — | — | |
| Llama-3.3-70B-InstructParameters=70B2026.01 | 92.9 | — | — | — | |
| Chain-of-Thought2026.02 | 92.8 | — | — | — | |
| SPP2026.02 | 92.8 | — | — | — | |
| Llama 70BShots=5-shot2026.02 | 92.65 | — | — | — | |
| Single2026.02 | 92.4 | — | — | — | |
| Gemma 2 27BShots=5-shot2026.02 | 92.34 | — | — | — | |
| MiMo-V2-Flash Base# Shots=8-shot, # Activated Params=15B, # Total Params=309B2026.01 | 92.3 | — | — | — | |
| Kimi-K2 Base# Shots=8-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 92.1 | — | — | — | |
| LLM-Debate2026.02 | 91.6 | — | — | — | |
| DeepSeek-V3.1 Base# Shots=8-shot, # Activated Params=37B, # Total Params=671B2026.01 | 91.4 | — | — | — | |
| SDAR 8BParameters=8B, Sampling Strategy=Greedy2025.12 | 91.3 | — | — | — | |
| EFFGENBackbone=Qwen2.5-7B-Instruct2026.01 | 91.28 | — | — | — | |
| DyLAN2026.02 | 91.2 | — | — | — | |
| DeepSeek-V3.2 Exp Base# Shots=8-shot, # Activated Params=37B, # Total Params=671B2026.01 | 91.1 | — | — | — | |
| NBDiff-7B-INSTRUCTParameters=7B, Training Protocol=Instruct, Sampling Strategy=Greedy2025.12 | 91 | — | — | — | |
| Self-Refine2026.02 | 90.8 | — | — | — | |
| Raw ModelBackbone=Qwen2.5-7B-Instruct2026.01 | 90.75 | — | — | — | |
| SmolagentsBackbone=Qwen2.5-14B-Instruct2026.01 | 90.45 | — | — | — | |
| FiMI InstructShots=5-shot2026.02 | 90.37 | — | — | — | |
| AutogenBackbone=Qwen2.5-7B-Instruct2026.01 | 90.3 | — | — | — | |
| LangChainBackbone=Qwen2.5-14B-Instruct2026.01 | 89.99 | — | — | — | |
| ProFitModel=Qwen3-14B-Base, Evaluation Samples=8 samples2026.01 | 89.62 | — | — | — | |
| In-Writing-BaseBackbone=Qwen3-8B, Shot=12026.01 | 89.5 | — | — | — | |
| Phi-42026.01 | 89.4 | — | — | — | |
| In-Writing-BaseBackbone=Qwen3-8B, Shot=42026.01 | 89.4 | — | — | — | |
| Mistral 24BShots=5-shot2026.02 | 89.23 | — | — | — | |
| PADsource_model=Qwen3-32B, training_dataset=gsm8k2026.02 | 89.05 | — | — | — | |
| LLaDA2.0-mini preview 16BA1BParameters=16BA1B, Sampling Strategy=Greedy2025.12 | 89 | — | — | — | |
| DFTModel=Qwen3-14B-Base, Evaluation Samples=8 samples2026.01 | 88.98 | — | — | — | |
| PADsource_model=Gemma3-27B-it, training_dataset=gsm8k2026.02 | 88.48 | — | — | — | |
| PADsource_model=Qwen3-8B, training_dataset=gsm8k2026.02 | 88.42 | — | — | — | |
| PPCVBackbone=Llama-3.1-8B-Instruct2026.02 | 88.24 | — | — | — | |
| Original Qwen3-8Bstatus=baseline2026.02 | 88.17 | — | — | — | |
| Full2025.12 | 88 | — | — | — | |
| DFTModel=Qwen3-4B-Base, Evaluation Samples=8 samples2026.01 | 87.83 | — | — | — | |
| DAsource_model=Gemma3-27B-it, training_dataset=gsm8k2026.02 | 87.79 | — | — | — | |
| ProFitModel=Qwen3-4B-Base, Evaluation Samples=8 samples2026.01 | 87.55 | — | — | — | |
| RSFTsource_model=Qwen3-8B, training_dataset=gsm8k2026.02 | 87.11 | — | — | — | |
| DAsource_model=Qwen3-32B, training_dataset=gsm8k2026.02 | 87.11 | — | — | — | |
| StreamingLLMKV cache budget=5122025.12 | 87 | — | — | — | |
| RSFTsource_model=Gemma3-27B-it, training_dataset=gsm8k2026.02 | 86.96 | — | — | — | |
| RSFTsource_model=Qwen3-32B, training_dataset=gsm8k2026.02 | 86.73 | — | — | — | |
| Phi-DecodingBackbone=Llama-3.1-8B-Instruct2026.02 | 86.58 | — | — | — | |
| DAsource_model=Qwen3-8B, training_dataset=gsm8k2026.02 | 86.43 | — | — | — | |
| In-Writing-BFBackbone=Qwen3-8B, Shot=12026.01 | 86 | — | — | — | |
| In-Writing-BFBackbone=Qwen3-8B, Shot=42026.01 | 86 | — | — | — | |
| In-Writing-IFBackbone=Qwen3-8B, Shot=12026.01 | 85.6 | — | — | — | |
| Llama3.1-8B-Instruct BaselineMP=false, bits=16, Backbone=Llama3.1-8B-Instruct2026.02 | 85.52 | — | — | — | |
| In-Writing-IFBackbone=Qwen3-8B, Shot=42026.01 | 85.4 | — | — | — | |
| DIRBackbone=Llama3.1-8B-Instruct2025.12 | 84.84 | — | — | — | |
| EFFGENBackbone=Qwen2.5-3B-Instruct2026.01 | 84.83 | — | — | — | |
| Foundation-Sec-8B-InstructParameters=8B2026.01 | 84.8 | — | — | — | |
| SKBackbone=Llama3.1-8B-Instruct2025.12 | 84.61 | — | — | — | |
| Reflective Confidence2025.12 | 84.6 | — | — | — | |
| LangChainBackbone=Qwen2.5-7B-Instruct2026.01 | 84.38 | — | — | — | |
| ALBMBackbone=Llama3.1-8B-Instruct2025.12 | 84.08 | — | — | — | |
| SmolagentsBackbone=Qwen2.5-7B-Instruct2026.01 | 84 | — | — | — | |
| StreamingLLMKV cache budget=3842025.12 | 84 | — | — | — | |
| BaseBackbone=Llama3.1-8B-Instruct2025.12 | 83.93 | — | — | — | |
| InfoRMBackbone=Llama3.1-8B-Instruct2025.12 | 83.78 | — | — | — | |
| PoEBackbone=Llama3.1-8B-Instruct2025.12 | 83.62 | — | — | — | |
| VanillaModel=Qwen3-14B-Base, Evaluation Samples=8 samples2026.01 | 83.44 | — | — | — |