General Knowledge on MMLU (Score)
90.1ScoreQwen 3 VL 32B Think
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 3 VL 32B ThinkModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=true2025.12 | 90.1 | |
| Qwen 3 32BModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=false2025.12 | 88.8 | |
| K2-V2 70B InstructModel Family=K2, Parameter Count=70B, Thinking Capability=false2025.12 | 88.4 | |
| DS-R1 32BModel Family=DeepSeek-R1, Parameter Count=32B, Thinking Capability=true2025.12 | 88 | |
| DeepSeek V3.2Model Variant=Exp Base, # Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.02 | 87.8 | |
| Kimi-K2Model Variant=Base, # Shots=5-shot, # Activated Params=32B, # Total Params=1043B2026.02 | 87.8 | |
| DeepSeek V3.1Model Variant=Base, # Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.02 | 87.4 | |
| MiMo-V2 FlashModel Variant=Base, # Shots=5-shot, # Activated Params=15B, # Total Params=309B2026.02 | 86.7 | |
| Qwen 3 VL 8B Think2025.12 | 86.5 | |
| Olmo 3.1 Think 32BTraining Stage=Final Think 3.1, Model Family=Olmo 3.1, Parameter Count=32B, Thinking Capability=true2025.12 | 86.4 | |
| GLM-4.5Model Variant=Base, # Shots=5-shot, # Activated Params=32B, # Total Params=355B2026.02 | 86.1 | |
| Step 3.5 FlashModel Variant=Base, # Shots=5-shot, # Activated Params=11B, # Total Params=196B2026.02 | 85.8 | |
| Olmo 3 Think (Final 3.0)Training Stage=Final Think 3.0, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 85.4 | |
| Qwen 3 8B2025.12 | 85.4 | |
| Olmo 3 Think (SFT)Training Stage=SFT, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 85.3 | |
| Olmo 3 Think (DPO)Training Stage=DPO, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 85.2 | |
| Nemotron Nano 9B v22025.12 | 84.3 | |
| OR Nemotron 7B2025.12 | 80.7 | |
| Olmo 3 7B ThinkStage=Final Think2025.12 | 77.8 | |
| OpenThinker3 7B2025.12 | 77.4 | |
| Olmo 3 7B ThinkStage=SFT2025.12 | 74.9 | |
| Olmo 3 7B ThinkStage=DPO2025.12 | 74.8 | |
| DS-R1 Qwen 7B2025.12 | 67.9 | |
| KEELArchitecture=512 Layers / 3B Params, Peak Learning Rate=4.5 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=5-shot2026.01 | 62.7 | |
| Pre-LNArchitecture=512 Layers / 3B Params, Peak Learning Rate=3.0 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=5-shot2026.01 | 59.5 |