Multitask Language Understanding on MMLU Pro (pass@1)
86.7pass@1Qwen 3.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 3.5Zero-shot=true, Parameters=122B-A10B, Organization=Alibaba Cloud2026.05 | 86.7 | |
| Nemotron 3 SuperZero-shot=true, Parameters=120B-A12B, Organization=Nvidia2026.05 | 83.73 | |
| Llama 4 MaverickZero-shot=true, Parameters=400B-A17B, Organization=MetaAI2026.05 | 80.62 | |
| Phoenix-VL 1.5 MediumZero-shot=true, Parameters=123B, Organization=HTX2026.05 | 76.81 | |
| GLM-4.5VZero-shot=true, Parameters=106B-A12B, Organization=Z.ai2026.05 | 72.17 | |
| UNA-score (MSE)Base Model=Qwen 8B, Alignment Strategy=UNA-score (MSE)2024.08 | 48.94 | |
| DPOBase Model=Qwen 8B, Alignment Strategy=DPO2024.08 | 47.48 | |
| UNA-pairwiseBase Model=Qwen 8B, Alignment Strategy=UNA-pairwise2024.08 | 47.48 | |
| Qwen 8BBase Model=Qwen 8B, Alignment Strategy=Baseline2024.08 | 47.21 | |
| KTOBase Model=Qwen 8B, Alignment Strategy=KTO2024.08 | 47.18 | |
| UNA-binary (BCE)Base Model=Qwen 8B, Alignment Strategy=UNA-binary (BCE)2024.08 | 46.89 | |
| UNA-score & binaryBase Model=Qwen 8B, Alignment Strategy=UNA-score & binary2024.08 | 44.83 | |
| UNA-score (MSE)Base Model=Llama 8B, Alignment Strategy=UNA-score (MSE)2024.08 | 34.42 | |
| UNA-score & binaryBase Model=Llama 8B, Alignment Strategy=UNA-score & binary2024.08 | 34.25 | |
| DPOBase Model=Llama 8B, Alignment Strategy=DPO2024.08 | 33.05 | |
| UNA-pairwiseBase Model=Llama 8B, Alignment Strategy=UNA-pairwise2024.08 | 33.05 | |
| UNA-binary (BCE)Base Model=Llama 8B, Alignment Strategy=UNA-binary (BCE)2024.08 | 33.01 | |
| KTOBase Model=Llama 8B, Alignment Strategy=KTO2024.08 | 32.86 | |
| Llama 8BBase Model=Llama 8B, Alignment Strategy=Baseline2024.08 | 32.73 | |
| UNA-binary (BCE)Base Model=Mistral 7B, Alignment Strategy=UNA-binary (BCE)2024.08 | 30.73 | |
| KTOBase Model=Mistral 7B, Alignment Strategy=KTO2024.08 | 30.43 | |
| DPOBase Model=Mistral 7B, Alignment Strategy=DPO2024.08 | 30.41 | |
| UNA-pairwiseBase Model=Mistral 7B, Alignment Strategy=UNA-pairwise2024.08 | 30.41 | |
| Mistral 7BBase Model=Mistral 7B, Alignment Strategy=Baseline2024.08 | 30.11 | |
| UNA-score & binaryBase Model=Mistral 7B, Alignment Strategy=UNA-score & binary2024.08 | 30.09 | |
| UNA-score (MSE)Base Model=Mistral 7B, Alignment Strategy=UNA-score (MSE)2024.08 | 29.72 | |
| UNA-score (MSE)Base Model=Gemma 4B, Alignment Strategy=UNA-score (MSE)2024.08 | 28.72 | |
| UNA-score & binaryBase Model=Gemma 4B, Alignment Strategy=UNA-score & binary2024.08 | 28.49 | |
| DPOBase Model=Gemma 4B, Alignment Strategy=DPO2024.08 | 28.03 | |
| UNA-pairwiseBase Model=Gemma 4B, Alignment Strategy=UNA-pairwise2024.08 | 28.03 | |
| KTOBase Model=Gemma 4B, Alignment Strategy=KTO2024.08 | 27.98 | |
| UNA-binary (BCE)Base Model=Gemma 4B, Alignment Strategy=UNA-binary (BCE)2024.08 | 27.95 | |
| Gemma 4BBase Model=Gemma 4B, Alignment Strategy=Baseline2024.08 | 27.92 | |
| InfLLM-v2Model Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 0.793 | |
| SPLAModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 0.793 | |
| SPAModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 0.791 | |
| Dense AttentionModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 0.789 | |
| NSAModel Size=14B, Temperature=0.6, Max output length=32k, Training context length=32k, Training tokens=1.2T, Sampling strategy=pass@12026.01 | 0.688 |