General Knowledge on MMLU-Pro (Accuracy)
86.7AccuracyQwen3.5-122B-A10B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-122B-A10B2026.04 | 86.7 | |
| Qwen-3.5 27BSpeedup @32k=0.55x, Speedup @16k=0.5x2026.04 | 86.1 | |
| SuperPrecision=BF162026.07 | 83.8 | |
| Nemotron 3 Super2026.04 | 83.73 | |
| SuperPrecision=NVFP42026.07 | 83.5 | |
| Puzzle-75B-A9BPrecision=BF162026.07 | 82.4 | |
| Puzzle-75B-A9BPrecision=NVFP42026.07 | 82.2 | |
| Qwen3.5-35B-A3B#Token=40002026.04 | 81.6 | |
| JoyAI-LLM Flash#Token=9002026.04 | 81.6 | |
| GPT-OSS-120B2026.04 | 81 | |
| Apriel-1.6Speedup @32k=1.0x, Speedup @16k=1.0x2026.04 | 78.8 | |
| Qwen3-30B-A3B#Token=16002026.04 | 78.7 | |
| Nemotron-3-Nano 30BSpeedup @32k=4.09x, Speedup @16k=2.8x2026.04 | 78.3 | |
| Nemotron-Nano 12B v2Speedup @32k=5.85x, Speedup @16k=4.3x2026.04 | 77.2 | |
| all-attentionSpeedup @32k=1.0x, Speedup @16k=1.0x2026.04 | 76.8 | |
| Qwen3-Next-80B-A3B#Token=15002026.04 | 76.7 | |
| Reg|Lklhd–18Speedup @32k=4.76x, Speedup @16k=2.2x2026.04 | 76.3 | |
| Idealized|All–18Speedup @32k=1.99x, Speedup @16k=1.1x2026.04 | 76.3 | |
| Reg|Lklhd–26Speedup @32k=2.85x, Speedup @16k=1.5x2026.04 | 76.2 | |
| GLM-4.7-Flash-T#Token=62002026.04 | 75.8 | |
| S1: Distil. Idealized|All–6Speedup @32k=6.13x, Speedup @16k=2.5x2026.04 | 74.5 | |
| Idealized|Lklhd–6Speedup @32k=6.2x, Speedup @16k=2.4x2026.04 | 73.6 | |
| Idealized|All–6Speedup @32k=6.13x, Speedup @16k=2.5x2026.04 | 73.3 | |
| JoyAI-LLM Flash-BaseEvaluation Framework=OpenCompass, Decoding=Greedy2026.04 | 73.1 | |
| QwenCheckpoint=SFTMaxOOD2025.09 | 73 | |
| Falcon-H1R 7BSpeedup @32k=4.61x, Speedup @16k=3.4x2026.04 | 72.5 | |
| Reg|Lklhd–13Speedup @32k=6.9x, Speedup @16k=2.7x2026.04 | 69.3 | |
| Reg|Lklhd–10Speedup @32k=10.69x, Speedup @16k=4.2x2026.04 | 68.2 | |
| Apriel-H1 15BSpeedup @32k=1.97x, Speedup @16k=1.9x2026.04 | 67.5 | |
| QwenCheckpoint=SFTEnd2025.09 | 66 | |
| QwenCheckpoint=RLEnd2025.09 | 66 | |
| OLMo-Hybrid-Think 7BSpeedup @32k=2.51x, Speedup @16k=2.1x2026.04 | 65.7 | |
| Qwen3-30B-A3B-BaseEvaluation Framework=OpenCompass, Decoding=Greedy2026.04 | 61.7 | |
| Qwen3.5-35B-A3B-BaseEvaluation Framework=OpenCompass, Decoding=Greedy2026.04 | 60.7 | |
| SDAR-8BScale=8B, Tokens=55B, FLOPs=2640, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 56.9 | |
| OPDLM-8BScale=8B, Tokens=0.066B, FLOPs=4.2, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 53.7 | |
| Qwen3Model Version=4B, Architecture=Dense, # Params=4.0B, # Tokens=36T2025.10 | 51.4 | |
| SDAR-4BScale=4B, Tokens=55B, FLOPs=1320, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 50.9 | |
| OuroModel Version=1.4B R4, Architecture=LoopLM, # Params=1.4B, # Tokens=7.7T, recurrent steps=42025.10 | 48.62 | |
| OPDLM-4BScale=4B, Tokens=0.076B, FLOPs=2.4, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 46.3 | |
| LLaMACheckpoint=SFTMaxOOD2025.09 | 46 | |
| Dream-7BScale=8B, Tokens=580B, FLOPs=24360, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 43.3 | |
| Fast-dLLM-v2-7BScale=8B, Tokens=1B, FLOPs=42, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 41.5 | |
| MSA-CPTTraining Strategy=sparse continued pretraining2026.06 | 39.1 | |
| MSA-PTTraining Strategy=from-scratch sparse pretraining2026.06 | 38.8 | |
| FullTraining Strategy=Full-Attention baseline2026.06 | 38.5 | |
| Qwen2.5Model Version=3B, Architecture=Dense, # Params=3.0B, # Tokens=18T2025.10 | 37.87 | |
| Qwen3Model Version=1.7B, Architecture=Dense, # Params=1.7B, # Tokens=36T2025.10 | 37.27 | |
| LLaMACheckpoint=RLEnd2025.09 | 37 | |
| LLaDA-8BScale=8B, Tokens=1500B, FLOPs=72000, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 37 | |
| LLaMACheckpoint=SFTEnd2025.09 | 36 | |
| Gemma3Model Version=4B, Architecture=Dense, # Params=4.0B, # Tokens=4T2025.10 | 34.61 | |
| Llama3.2Model Version=3B, Architecture=Dense, # Params=3.0B, # Tokens=9T2025.10 | 33.34 | |
| Qwen2.5Model Version=1.5B, Architecture=Dense, # Params=1.5B, # Tokens=18T2025.10 | 29.11 | |
| EMM2026.06 | 28.1 | |
| MERGEvolveProtocol=Cross-domain generalization2026.06 | 27.17 | |
| Pack of LLMs2026.06 | 27.07 | |
| LoraHubProtocol=Cross-domain generalization2026.06 | 27.03 | |
| Pack of LLMsProtocol=Cross-domain generalization2026.06 | 26.97 | |
| Model SwarmsProtocol=Cross-domain generalization2026.06 | 26.92 | |
| MERGEvolve2026.06 | 26.63 | |
| EMMProtocol=Cross-domain generalization2026.06 | 26.54 | |
| Model Swarms2026.06 | 26.47 | |
| LoraHub2026.06 | 26.15 | |
| Expert Fusion2026.06 | 25.3 | |
| Llama3.2Model Version=1.2B, Architecture=Dense, # Params=1.0B, # Tokens=9T2025.10 | 11.8 | |
| Gemma3Model Version=1B, Architecture=Dense, # Params=1.0B, # Tokens=2T2025.10 | 11.31 |