Massive Multitask Language Understanding on MMLU (FL/PM/All Scores)
61.9MMLU FL ScoreHD + Confidence
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HD + ConfidenceModel=Qwen-2.5-7B-Instruct2026.01 | 61.9 | 82.5 | 76.9 | |
| HD + Learn2AggModel=Qwen-2.5-7B-Instruct2026.01 | 60.3 | 80.5 | 75.2 | |
| High DiversityModel=Qwen-2.5-7B-Instruct2026.01 | 58.7 | 82.2 | 74.3 | |
| ConfidenceModel=Qwen-2.5-7B-Instruct2026.01 | 58.7 | 81.8 | 76.9 | |
| Debate 5 × 5Model=Qwen-2.5-7B-Instruct2026.01 | 57.1 | 73.4 | 72.7 | |
| HD + ConfidenceModel=Llama-3.1-8B-Instruct2026.01 | 55.5 | 84.2 | 71.3 | |
| HD + Learn2AggModel=Llama-3.1-8B-Instruct2026.01 | 55.1 | 79.7 | 70 | |
| Majority VoteModel=Qwen-2.5-7B-Instruct2026.01 | 54.8 | 80.6 | 76.4 | |
| High DiversityModel=Llama-3.1-8B-Instruct2026.01 | 54.8 | 80.5 | 68.7 | |
| ConfidenceModel=Llama-3.1-8B-Instruct2026.01 | 54.8 | 81.8 | 72 | |
| Single ModelModel=Qwen-2.5-7B-Instruct2026.01 | 54.7 | 76.5 | 73.5 | |
| Debate 5 × 5Model=Llama-3.1-8B-Instruct2026.01 | 54.4 | 77.2 | 68 | |
| Majority VoteModel=Llama-3.1-8B-Instruct2026.01 | 53.2 | 80.9 | 71.1 | |
| Single ModelModel=Llama-3.1-8B-Instruct2026.01 | 50.3 | 79.8 | 70.1 |