Language Understanding on MMLU
82.33RACortexDebate
Evaluation Results
| Method | Links | |
|---|---|---|
| CortexDebateType=Ours2025.07 | 82.33 | |
| PRDType=Full Debate2025.07 | 77.33 | |
| RECONCILEType=Full Debate2025.07 | 75 | |
| GDType=Part Debate2025.07 | 74 | |
| ChatEvalType=Full Debate2025.07 | 73 | |
| NDType=Part Debate2025.07 | 71.67 | |
| MLDType=Full Debate2025.07 | 71.33 | |
| MaVType=No Debate2025.07 | 69.33 | |
| LLAMA 2Size=70B, Shots=5-shot2023.07 | 68.9 | |
| LLAMA 1Size=65B, Shots=5-shot2023.07 | 63.4 | |
| LLAMA 2Size=34B, Shots=5-shot2023.07 | 62.6 | |
| KEELTraining=SFT, Evaluation Setting=5-shot2026.01 | 62.5 | |
| Pre-LNTraining=SFT, Evaluation Setting=5-shot2026.01 | 60 | |
| Qwen-7BModel Size=7B, Status=final released2023.08 | 58.2 | |
| LLAMA 1Size=33B, Shots=5-shot2023.07 | 57.8 | |
| FalconSize=40B, Shots=5-shot2023.07 | 55.4 | |
| LLAMA 2Size=13B, Shots=5-shot2023.07 | 54.8 | |
| Baichuan2-7BModel Size=7B2023.08 | 54.2 | |
| InternLM-7BModel Size=7B2023.08 | 51 | |
| Qwen-VLInitialization=Intermediate Qwen-7B checkpoint2023.08 | 50.7 | |
| Qwen-7BModel Size=7B, Status=intermediate, Usage=LLM initialization for Qwen-VL2023.08 | 49.9 | |
| ChatGLM2-6BModel Size=6B2023.08 | 47.9 | |
| MPTSize=30B, Shots=5-shot2023.07 | 46.9 | |
| LLAMA 1Size=13B, Shots=5-shot2023.07 | 46.9 | |
| LLAMA2-7BModel Size=7B2023.08 | 46.8 | |
| LLAMA 2Size=7B, Shots=5-shot2023.07 | 45.3 | |
| Baichuan-7BModel Size=7B2023.08 | 42.3 | |
| LLAMA 1Size=7B, Shots=5-shot2023.07 | 35.1 | |
| LLAMA-7BModel Size=7B2023.08 | 35.1 | |
| MPTSize=7B, Shots=5-shot2023.07 | 26.8 | |
| FalconSize=7B, Shots=5-shot2023.07 | 26.2 |