Knowledge on MMLU
85.93AccuracyDictaLM 3.0 24B-Think
Evaluation Results
| Method | Links | |
|---|---|---|
| DictaLM 3.0 24B-ThinkParameters=24B, Variant=Thinking2026.02 | 85.93 | |
| Llama-3.3-70B-InstructShots=52026.04 | 85.86 | |
| Llama-3.3-70B-InstructParameters=70B2026.01 | 85.7 | |
| Phi-42026.01 | 84.9 | |
| Qwen3.5-9BShots=52026.04 | 84.58 | |
| Mistral Small 3.12026.02 | 83.11 | |
| ReasonAnyModel Family=Qwen2.5-32B2026.01 | 82.09 | |
| Qwen3-14BShots=52026.04 | 81.94 | |
| LinearModel Family=Qwen2.5-32B2026.01 | 81.9 | |
| Task ArithmeticModel Family=Qwen2.5-32B2026.01 | 81.89 | |
| DAREModel Family=Qwen2.5-32B2026.01 | 81.85 | |
| Qwen2.5-32B-Instruct (Safety)Model Family=Qwen2.5-32B2026.01 | 81.78 | |
| LEDModel Family=Qwen2.5-32B2026.01 | 80.83 | |
| FuseLLMModel Family=Qwen2.5-32B2026.01 | 80.62 | |
| DeepSeek-R1-Distill-Qwen-32B (Reasoning)Model Family=Qwen2.5-32B2026.01 | 79.65 | |
| TIESModel Family=Qwen2.5-32B2026.01 | 79.52 | |
| ReasonAnyModel Family=Qwen2.5-14B2026.01 | 78.86 | |
| Task ArithmeticModel Family=Qwen2.5-14B2026.01 | 78.79 | |
| Qwen2.5-14B-Instruct (Safety)Model Family=Qwen2.5-14B2026.01 | 78.73 | |
| XekRung-8BShots=52026.04 | 78.58 | |
| Qwen3-8BShots=52026.04 | 78.23 | |
| LinearModel Family=Qwen2.5-14B2026.01 | 77.63 | |
| LEDModel Family=Qwen2.5-14B2026.01 | 77.56 | |
| DAREModel Family=Qwen2.5-14B2026.01 | 76.94 | |
| SimPO2025.09 | 76.79 | |
| FuseLLMModel Family=Qwen2.5-14B2026.01 | 76.25 | |
| DPO2025.09 | 75.77 | |
| TD-MNPO2025.09 | 75.63 | |
| HT-MNPOReward Model=ArmoRM-Llama32025.09 | 75.6 | |
| HT-MNPOReward Model=Athene-RM-8B2025.09 | 75.4 | |
| HT-MNPOReward Model=Skywork-Reward-V22025.09 | 75.39 | |
| SPPO2025.09 | 75.37 | |
| SFT Model2025.09 | 75.35 | |
| INPO2025.09 | 74.79 | |
| Gemma 3 27BParameters=27B2026.02 | 74.6 | |
| SecGPT-14BShots=52026.04 | 73.73 | |
| DeepSeek-R1-Distill-Qwen-14B (Reasoning)Model Family=Qwen2.5-14B2026.01 | 73.28 | |
| TIESModel Family=Qwen2.5-14B2026.01 | 72.93 | |
| Qwen3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=36T2025.11 | 72.25 | |
| Llama-3.1-8B-InstructParameters=8B2026.01 | 69.8 | |
| Llama-3.1-8B-InstructShots=52026.04 | 69.4 | |
| Foundation-Sec-8B-ReasoningParameters=8B2026.01 | 68.3 | |
| Foundation-Sec-8B-ReasoningShots=52026.04 | 68.3 | |
| Llama-Primus-Merged2026.01 | 68.2 | |
| DSRMidtraining Schedule=DSR2026.05 | 67.27 | |
| ReasonAnyMerging Protocol=ReasonAny2026.01 | 67.15 | |
| SafetyBase Model=Llama-3.1-8B2026.01 | 67.14 | |
| Granite-4.0-HTotal Parameters=7B, Active Parameters=1B, Trained Tokens=15T2025.11 | 66.79 | |
| WSDMidtraining Schedule=WSD2026.05 | 66.58 | |
| Unified 15B MoEMidtraining Corpus=1k→8k2026.05 | 66.4 | |
| Foundation-Sec-8B-InstructParameters=8B2026.01 | 66 | |
| Cosine-decayMidtraining Schedule=Cos2026.05 | 65.88 | |
| Unified 15B MoEMidtraining Corpus=4k→8k2026.05 | 65.35 | |
| Task ArithmeticMerging Protocol=Task Arithmetic2026.01 | 64.93 | |
| Unified 15B MoEMidtraining Corpus=4k2026.05 | 64.9 | |
| LFM2-8B-A1BTotal Parameters=8.3B, Active Parameters=1.5B, Trained Tokens=13T2025.11 | 64.84 | |
| LFM2-2.6BTotal Parameters=2.6B, Active Parameters=2.6B, Trained Tokens=11T2025.11 | 64.42 | |
| LEDMerging Protocol=LED2026.01 | 63.45 | |
| LinearMerging Protocol=Linear2026.01 | 63.25 | |
| UM-190kBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot, Training Mix=UltraMix2025.11 | 62.79 | |
| UM-187kBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot, Training Mix=UltraMix2025.11 | 62.25 | |
| BaseBackbone=Llama3-8B, Unlearning Framework=FT & Model Editing2026.03 | 62.1 | |
| UM-170kBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot, Training Mix=UltraMix2025.11 | 61.79 | |
| DAREMerging Protocol=DARE2026.01 | 61.76 | |
| TuluDPOBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 61.11 | |
| ORPOBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 60.92 | |
| SFTBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 60.87 | |
| UltraFBBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 60.37 | |
| Llama-3.2-3BTotal Parameters=3.2B, Active Parameters=3.2B, Trained Tokens=9T2025.11 | 60.35 | |
| HelpSteerBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 59.85 | |
| SmolLM3-3BTotal Parameters=3.1B, Active Parameters=3.1B, Trained Tokens=11T2025.11 | 59.84 | |
| FuseLLMMerging Protocol=FuseLLM2026.01 | 59.6 | |
| MobileMoE-LActive Parameters=922M, Total Parameters=5.3B, Few-shot count=5-shot2026.05 | 59.6 | |
| Qwen3-1.7B# Total Params=1.7B, # Trained Tokens=36T2025.11 | 59.11 | |
| BaseBackbone=Zephyr-7B, Unlearning Framework=FT & Model Editing2026.03 | 58.5 | |
| Gemma-3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=4T2025.11 | 58.35 | |
| NPO_GDBackbone=Zephyr-7B, Unlearning Framework=FT & Model Editing2026.03 | 58 | |
| CodePrefBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 57.94 | |
| ALTERBackbone=Llama3-8B, Unlearning Framework=LoRA Variations(r=8)2026.03 | 57.8 | |
| RMUBackbone=Zephyr-7B, Unlearning Framework=FT & Model Editing2026.03 | 57.8 | |
| NPOBackbone=Zephyr-7B, Unlearning Framework=FT & Model Editing2026.03 | 57.8 | |
| NPO_KLBackbone=Zephyr-7B, Unlearning Framework=FT & Model Editing2026.03 | 57.4 | |
| ELMBackbone=Llama3-8B, Unlearning Framework=FT & Model Editing2026.03 | 57.2 | |
| DPOBackbone Model=Gemma2-2B-it2026.01 | 57.02 | |
| MobileMoE-LMem (GB)=2.75, INT4 QAT=true, 5-shot=true2026.05 | 57 | |
| GANPOBackbone Model=Gemma2-2B-it2026.01 | 56.93 | |
| TIESMerging Protocol=TIES2026.01 | 56.74 | |
| BaseBackbone Model=Gemma2-2B-it2026.01 | 56.73 | |
| ALTERBackbone=Zephyr-7B, Unlearning Framework=LoRA Variations(r=8)2026.03 | 56.4 | |
| ELMBackbone=Zephyr-7B, Unlearning Framework=FT & Model Editing2026.03 | 56.2 | |
| Unified 15B MoEPretraining Corpus=4k→8k2026.05 | 56.05 | |
| NPO_KLBackbone=Llama3-8B, Unlearning Framework=FT & Model Editing2026.03 | 56 | |
| AsymLoRABackbone=Llama3-8B, Unlearning Framework=LoRA Variations(r=8)2026.03 | 55.3 | |
| LFM2-1.2B# Total Params=1.2B, # Trained Tokens=11T2025.11 | 55.23 | |
| Unified 15B MoEPretraining Corpus=1k→8k2026.05 | 54.92 | |
| MobileMoE-MActive Parameters=528M, Total Parameters=2.8B, Few-shot count=5-shot2026.05 | 54.7 | |
| Unified 15B MoEPretraining Corpus=4k2026.05 | 54.29 | |
| AsymLoRABackbone=Zephyr-7B, Unlearning Framework=LoRA Variations(r=8)2026.03 | 54.1 | |
| ReasoningBase Model=Llama-3.1-8B2026.01 | 53.25 | |
| MobileMoE-MMem (GB)=1.48, INT4 QAT=true, 5-shot=true2026.05 | 52.4 |