General Reasoning on AGI Eval English
90.1ScoreQwen 3 VL 8B Think
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 3 VL 8B Think2025.12 | 90.1 | |
| Qwen 3 8B2025.12 | 87 | |
| Qwen 3 VL 8B Inststage=Instruct2025.12 | 84.5 | |
| Nemotron Nano 9B v22025.12 | 83.1 | |
| Olmo 3 7B ThinkStage=Final Think2025.12 | 81.5 | |
| OR Nemotron 7B2025.12 | 81.4 | |
| Olmo 3 7B ThinkStage=DPO2025.12 | 79.1 | |
| OpenThinker3 7B2025.12 | 78.6 | |
| Olmo 3 7B ThinkStage=SFT2025.12 | 77.2 | |
| Qwen 3 8Bstage=Instruct2025.12 | 76 | |
| Qwen 2.5 7Bstage=Instruct2025.12 | 69.8 | |
| DS-R1 Qwen 7B2025.12 | 69.5 | |
| Olmo 3 7B Instructstage=Final Instruct2025.12 | 64.4 | |
| Olmo 3 7B Instructstage=DPO2025.12 | 64 | |
| Granite 3.3 8B Inststage=Instruct2025.12 | 64 | |
| Olmo 3 7B Instructstage=SFT2025.12 | 59.2 | |
| OLMo 2 7B Inststage=Instruct2025.12 | 56.1 | |
| LLAMA 2Size=70B, Shots=3-5 shot2023.07 | 54.2 | |
| Apertus 8B Inststage=Instruct2025.12 | 50.8 | |
| LLAMA 1Size=65B, Shots=3-5 shot2023.07 | 47.6 | |
| KEELArchitecture=512 Layers / 3B Params, Peak Learning Rate=4.5 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=0-shot2026.01 | 46.5 | |
| LLAMA 2Size=34B, Shots=3-5 shot2023.07 | 43.4 | |
| LLAMA 1Size=33B, Shots=3-5 shot2023.07 | 41.7 | |
| LLAMA 2Size=13B, Shots=3-5 shot2023.07 | 39.1 | |
| Pre-LNArchitecture=512 Layers / 3B Params, Peak Learning Rate=3.0 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=0-shot2026.01 | 37.9 | |
| FalconSize=40B, Shots=3-5 shot2023.07 | 37 | |
| LLAMA 1Size=13B, Shots=3-5 shot2023.07 | 33.9 | |
| MPTSize=30B, Shots=3-5 shot2023.07 | 33.8 | |
| LLAMA 2Size=7B, Shots=3-5 shot2023.07 | 29.3 | |
| LLAMA 1Size=7B, Shots=3-5 shot2023.07 | 23.9 | |
| MPTSize=7B, Shots=3-5 shot2023.07 | 23.5 | |
| FalconSize=7B, Shots=3-5 shot2023.07 | 21.2 |