Natural Language Understanding on AGIEval
71.6AccuracyLlama 3 405B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Llama 3 405BModel Type=Pre-trained2024.07 | 71.6 | — | — | |
| Llama 3 70BModel Type=Pre-trained2024.07 | 64.6 | — | — | |
| Mixtral 8x22BModel Type=Pre-trained2024.07 | 61.5 | — | — | |
| RLVRBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 58.1 | 53.4 | 37.7 | |
| C3RL w/o RefBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 58.1 | 68.2 | 19.7 | |
| C3RLBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 57.8 | 68.8 | 15.4 | |
| RLCRBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 57.5 | 66.4 | 20.5 | |
| SCBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 53.7 | 71.1 | 20.8 | |
| BaseBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 51 | 54.3 | 37.5 | |
| SaySelfBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 50.8 | 74.5 | 14.7 | |
| SFT+RefBackbone Model=Qwen2.5VL-7B-Instruct2026.07 | 49.1 | 75.3 | 10.5 | |
| Qwen2.52025.12 | 48.13 | — | — | |
| Llama 3 8BModel Type=Pre-trained2024.07 | 47.8 | — | — | |
| Gemma 7BModel Type=Pre-trained2024.07 | 46 | — | — | |
| Mistral 7BModel Type=Pre-trained2024.07 | 42.7 | — | — | |
| Qwen32025.12 | 42.14 | — | — | |
| RLVRBackbone Model=Llama-3.2-3B-Instruct2026.07 | 37.3 | 52.2 | 46.8 | |
| C3RLBackbone Model=Llama-3.2-3B-Instruct2026.07 | 37.3 | 66 | 24.6 | |
| C3RL w/o RefBackbone Model=Llama-3.2-3B-Instruct2026.07 | 37.2 | 65.2 | 21.8 | |
| NITPModel Scale=3B2026.05 | 35.11 | — | — | |
| BaseBackbone Model=Llama-3.2-3B-Instruct2026.07 | 34.8 | 55.7 | 42.3 | |
| RLCRBackbone Model=Llama-3.2-3B-Instruct2026.07 | 34.4 | 65.8 | 11.1 | |
| Gamayun2025.12 | 33.34 | — | — | |
| mPLUG-Owl2Shot setting=0-shot2023.11 | 32.7 | — | — | |
| SFT+RefBackbone Model=Llama-3.2-3B-Instruct2026.07 | 32.1 | 59.4 | 24 | |
| NTPModel Scale=3B2026.05 | 31.77 | — | — | |
| NITPModel Scale=2B2026.05 | 31.66 | — | — | |
| NTPModel Scale=2B2026.05 | 30.51 | — | — | |
| Arcanazero-shot=true2024.10 | 29.3 | — | — | |
| SCBackbone Model=Llama-3.2-3B-Instruct2026.07 | 28.6 | 64.4 | 35.2 | |
| LLaMA-2-ChatShot setting=0-shot2023.11 | 28.5 | — | — | |
| LLaMA-2-Chatzero-shot=true2024.10 | 28.5 | — | — | |
| Gemma32025.12 | 28.05 | — | — | |
| DeepSeek-VLVersion=7B Chat, Encoder=SigLIP+SAM, Evaluation Protocol=Generation-based (greedy decoding)2024.03 | 27.8 | — | — | |
| NTPModel Scale=0.5B2026.05 | 27.01 | — | — | |
| Llama3.22025.12 | 26.32 | — | — | |
| NITPModel Scale=0.5B2026.05 | 26.15 | — | — | |
| WizardLMShot setting=0-shot2023.11 | 23.2 | — | — | |
| WizardLMzero-shot=true2024.10 | 23.2 | — | — | |
| LLaMA-2Shot setting=0-shot2023.11 | 21.8 | — | — | |
| LLaMA-2zero-shot=true2024.10 | 21.8 | — | — | |
| Vicuna-v1.5Shot setting=0-shot2023.11 | 21.2 | — | — | |
| Vicuna-v1.5zero-shot=true2024.10 | 21.2 | — | — | |
| SaySelfBackbone Model=Llama-3.2-3B-Instruct2026.07 | 19.4 | 61.1 | 7.2 | |
| DeepSeek-LLMVersion=7B Chat, Encoder=None, Evaluation Protocol=Generation-based (greedy decoding)2024.03 | 19.3 | — | — | |
| DeepSeek-VLVersion=1B Chat, Encoder=SigLIP, Evaluation Protocol=Generation-based (greedy decoding)2024.03 | 14 | — | — |