Accuracy on MMLU-Pro
82.05AccuracySMCS
Evaluation Results
| Method | Links | |
|---|---|---|
| SMCS2025.07 | 82.05 | |
| GPT-4.12025.07 | 80.41 | |
| QwQ-32B2025.07 | 74.69 | |
| Self-MoA2025.07 | 69.76 | |
| HuggingGPTImplementation Framework=original2026.03 | 65.59 | |
| AgentVerseImplementation Framework=original2026.03 | 64.64 | |
| AgentVerseImplementation Framework=MASFactory2026.03 | 64.16 | |
| HuggingGPTImplementation Framework=MASFactory2026.03 | 63.66 | |
| CAMELImplementation Framework=MASFactory2026.03 | 63.04 | |
| OpenReasoner-Zero*Method Category=Previous RLVR Methods, Result Source=Reported from original papers2025.10 | 58.7 | |
| DIVERBase Model=DeepSeek-R1-Distill-Qwen-7B2025.09 | 56.6 | |
| GRPOBase Model=Qwen2.5-7B-Base, Reward Strategy=Clip-higher2025.09 | 55.4 | |
| DIVERBase Model=Qwen2.5-7B-Base2025.09 | 55.2 | |
| GRPOBase Model=DeepSeek-R1-Distill-Qwen-7B, Reward Strategy=Clip-higher2025.09 | 54.9 | |
| LUFFYMethod Category=Comparable Baseline Methods2025.10 | 53 | |
| Prefix-RFTMethod Category=Comparable Baseline Methods2025.10 | 52.1 | |
| DIVERBase Model=LLaMA-3.1-8B-Instruct2025.09 | 52 | |
| ReLIFTMethod Category=Comparable Baseline Methods2025.10 | 51.9 | |
| Vibe Graphing-Task SpecificImplementation Framework=Vibe Graphing2026.03 | 51.73 | |
| GRPOBase Model=LLaMA-3.1-8B-Instruct, Reward Strategy=Clip-higher2025.09 | 50.8 | |
| CAMELImplementation Framework=original2026.03 | 50.08 | |
| SFT+RLMethod Category=SFT and RL2025.10 | 49.6 | |
| Dream-Instruct 7B + G-StarRefinement Strategy=Loop-based refinement2025.10 | 47.9 | |
| ICPO†Method Category=Comparable Baseline Methods2025.10 | 47.6 | |
| Dream-Instruct 7BSource=Reproduced2025.10 | 46.9 | |
| ICPOMethod Category=Comparable Baseline Methods2025.10 | 46.5 | |
| SFTMethod Category=SFT and RL2025.10 | 44.4 | |
| Dream-Instruct 7BSource=Originally Published2025.10 | 43.3 | |
| Falcon3-7BEvaluation Protocol=Single, Backbone Model=Falcon3-7B2025.10 | 42.71 | |
| RLMethod Category=SFT and RL2025.10 | 42.6 | |
| M_RLP + PostPre-training Condition=Reinforcement Learning from Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=Standard2025.09 | 42.4 | |
| K-LVR (MCV)Evaluation Protocol=Ensemble (PoE), Vocabulary Strategy=MCV2025.10 | 42 | |
| Oat-Zero*Method Category=Previous RLVR Methods, Result Source=Reported from original papers2025.10 | 41.7 | |
| M_RLP + PostPre-training Condition=Reinforcement Learning from Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=@1[4]2025.09 | 41.3 | |
| K-LVR (1-Bytes)Evaluation Protocol=Ensemble (PoE), Vocabulary Strategy=1-Bytes2025.10 | 41.21 | |
| ReConcileBase Model=Gemma-3-4B2025.08 | 40 | |
| M_CPT + PostPre-training Condition=Continuous Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=Standard2025.09 | 39.92 | |
| M_CPT + PostPre-training Condition=Continuous Pretraining, Post-training=SFT + RLVR, Evaluation Protocol=@1[4]2025.09 | 38.49 | |
| M_base + PostPre-training Condition=Base, Post-training=SFT + RLVR, Evaluation Protocol=Standard2025.09 | 37.85 | |
| Qwen2.5-3BEvaluation Protocol=Single, Backbone Model=Qwen2.5-3B2025.10 | 37.71 | |
| M_base + PostPre-training Condition=Base, Post-training=SFT + RLVR, Evaluation Protocol=@1[4]2025.09 | 36.53 | |
| M_RLPPre-training Condition=Reinforcement Learning from Pretraining, Post-training=None, Evaluation Protocol=Standard2025.09 | 34.62 | |
| SimpleRL-Zero*Method Category=Previous RLVR Methods, Result Source=Reported from original papers2025.10 | 34.5 | |
| PRIME-Zero*Method Category=Previous RLVR Methods, Result Source=Reported from original papers2025.10 | 32.7 | |
| Single-Agent BestBase Model=Gemma-3-4B2025.08 | 32.3 | |
| DIVERBase Model=Qwen2.5-Math-1.5B2025.09 | 31.8 | |
| TMA-AllComponBase Model=Gemma-3-4B2025.08 | 31.7 | |
| M_RLPPre-training Condition=Reinforcement Learning from Pretraining, Post-training=None, Evaluation Protocol=@1[4]2025.09 | 30.8 | |
| MDAgentsBase Model=Gemma-3-4B2025.08 | 30.7 | |
| GRPOBase Model=Qwen2.5-Math-1.5B, Reward Strategy=Clip-higher2025.09 | 30.2 | |
| DyLanBase Model=Gemma-3-4B2025.08 | 28.7 | |
| M_basePre-training Condition=Base, Post-training=None, Evaluation Protocol=Standard2025.09 | 28.17 | |
| M_CPTPre-training Condition=Continuous Pretraining, Post-training=None, Evaluation Protocol=Standard2025.09 | 27.81 | |
| MedAgentsBase Model=Gemma-3-4B2025.08 | 24.7 | |
| M_CPTPre-training Condition=Continuous Pretraining, Post-training=None, Evaluation Protocol=@1[4]2025.09 | 24.61 | |
| M_basePre-training Condition=Base, Post-training=None, Evaluation Protocol=@1[4]2025.09 | 23.95 | |
| COACTTraining Dataset=MATH2026.04 | 23.42 | |
| Pref + EntTraining Dataset=MATH2026.04 | 22.76 | |
| EntropyTraining Dataset=MATH2026.04 | 22.15 | |
| Pref CertaintyTraining Dataset=MATH2026.04 | 22.08 | |
| RandomTraining Dataset=MATH2026.04 | 21.34 | |
| Qwen2.5-Math-7BMethod Category=Base Model2025.10 | 16.9 | |
| Naive (MCV)Evaluation Protocol=Ensemble (PoE), Vocabulary Strategy=Naive (MCV)2025.10 | 2.07 | |
| UnionEvaluation Protocol=Ensemble (PoE), Vocabulary Strategy=Union2025.10 | 1.93 |