Question Answering on MMLU
88.7AccuracyQwen 3 VL 32B Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 3 VL 32B InstructParameters=32B2025.12 | 88.7 | |
| GPT-4Number of shots=5-shot2023.11 | 86.5 | |
| Qwen 3 32BThinking=No, Parameters=32B2025.12 | 85.8 | |
| Qwen 2.5 32BParameters=32B2025.12 | 84.6 | |
| Qwen 3 VL 8B Inststage=Instruct2025.12 | 83.6 | |
| Olmo 3.1 32B InstructStage=DPO2025.12 | 81.9 | |
| Olmo 3.1 32B InstructStage=Final Instruct 3.12025.12 | 80.9 | |
| Qwen 3 8Bstage=Instruct2025.12 | 80.4 | |
| Olmo 3.1 32B InstructStage=SFT2025.12 | 79 | |
| Qwen 2.5 7Bstage=Instruct2025.12 | 77.2 | |
| OLMo 2 32BParameters=32B2025.12 | 77.1 | |
| Gemma 2 27BParameters=27B2025.12 | 76.1 | |
| Gemma 3 27BParameters=27B2025.12 | 74.6 | |
| FalconModel Size=180B2023.11 | 70.6 | |
| Falcon-180BNumber of shots=5-shot2023.11 | 70.6 | |
| Mixtral 8x7B2024.01 | 70.6 | |
| Apertus 70BParameters=70B2025.12 | 70.2 | |
| GPT-3.5Number of shots=5-shot2023.11 | 70 | |
| GPT-3.52024.01 | 70 | |
| LLaMA 2 70B2024.01 | 69.9 | |
| PaLM2023.11 | 69.3 | |
| Olmo 3 7B Instructstage=DPO2025.12 | 69.1 | |
| Olmo 3 7B Instructstage=Final Instruct2025.12 | 69.1 | |
| ExpGraphBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 69 | |
| LLaMA-2Model Size=70B2023.11 | 68.9 | |
| RCOBase Model=LLaMA-3-8B-Instruct2025.06 | 67.74 | |
| LLaMA-3-70B-Instruct2025.06 | 67.39 | |
| Olmo 3 7B Instructstage=SFT2025.12 | 67.1 | |
| LLaDA-8B-InstructSettings=Official scoreboard evaluation settings2026.01 | 65.4 | |
| CE OnlyTraining=CE Only (Matched Baseline)2026.01 | 65.1 | |
| RCOBase Model=UltraCM-13B2025.06 | 64.42 | |
| CTC+CETraining=CTC-trained2026.01 | 64.1 | |
| Granite 3.3 8B Inststage=Instruct2025.12 | 63.5 | |
| RCOBase Model=LLaMA-2-7B-Chat2025.06 | 63.17 | |
| Apertus 8B Inststage=Instruct2025.12 | 62.7 | |
| LLaMA-2Model Size=34B2023.11 | 62.6 | |
| RCOBase Model=LLaMA-2-13B-Chat2025.06 | 62.48 | |
| OLMo 2 7B Inststage=Instruct2025.12 | 61.6 | |
| DPCOBase Model=LLaMA-2-7B-Chat2025.06 | 60.46 | |
| RCOBase Model=Auto-J-13B2025.06 | 60.37 | |
| DPCOBase Model=UltraCM-13B2025.06 | 60.03 | |
| LLAMA-2-70B-Chat2025.06 | 60.02 | |
| ExpGraphBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 60 | |
| DPCOBase Model=LLaMA-2-13B-Chat2025.06 | 59.42 | |
| LLaMA-3-8B-Instruct2025.06 | 59.41 | |
| DPCOBase Model=LLaMA-3-8B-Instruct2025.06 | 59.21 | |
| DPCOBase Model=Auto-J-13B2025.06 | 58.47 | |
| LLaMA-2-13B-Chat2025.06 | 57.33 | |
| UltraCM-13B2025.06 | 57.19 | |
| FalconModel Size=40B2023.11 | 57 | |
| LLaMA-2-7B-Chat2025.06 | 56.96 | |
| STABLEPROMPTPrompt LM=Gemma-7B, Target LM=Gemma-7B2026.02 | 55.8 | |
| GFLOWPOPrompt LM=Gemma-7B, Target LM=Gemma-7B2026.02 | 55.6 | |
| LLaMA-2Model Size=13B2023.11 | 54.8 | |
| Self-refinement2025.06 | 54.79 | |
| ProTEGIPrompt LM=Gemma-7B, Target LM=Gemma-7B2026.02 | 53.5 | |
| Auto-J-13B2025.06 | 52.9 | |
| Aligner2025.06 | 52.56 | |
| Initial Answer2025.06 | 52.48 | |
| APEPrompt LM=Gemma-7B, Target LM=Gemma-7B2026.02 | 52.1 | |
| LLaMA-2Model Size=7B2023.11 | 45.3 | |
| OLTQAFull Config=True2023.05 | 36.14 | |
| UnifiedQA2023.05 | 28.77 | |
| FalconModel Size=7B2023.11 | 28 | |
| DCDM (MoE)Scale=1.5B, Active Params=1.2B (2.8B), In-context examples=0-shot2026.05 | 26.1 | |
| DCDMScale=1.5B, Active Params=1.5B, In-context examples=0-shot2026.05 | 25.59 | |
| BDLMScale=1.5B, Active Params=1.5B, In-context examples=0-shot2026.05 | 25.22 | |
| DCDM (MoE)Scale=0.5B, Active Params=0.4B (0.8B), In-context examples=0-shot2026.05 | 25.09 | |
| MDLMScale=1.5B, Active Params=1.5B, In-context examples=0-shot2026.05 | 24.49 | |
| DCDMScale=0.5B, Active Params=0.5B, In-context examples=0-shot2026.05 | 24.32 | |
| BDLMScale=0.5B, Active Params=0.5B, In-context examples=0-shot2026.05 | 23.53 | |
| MDLMScale=0.5B, Active Params=0.5B, In-context examples=0-shot2026.05 | 23.24 | |
| Softmaxzero-shot=true2025.01 | 22.97 | |
| LSSARzero-shot=true, p=152025.01 | 22.97 |