Multiple-choice Question Answering on KNIGHT Average
0.9392AccuracyHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| Humann=2002026.02 | 0.9392 | |
| GPT-4o2026.02 | 0.9052 | |
| Mistral Large2026.02 | 0.8959 | |
| Llama3-70B-InstructSize=70B, Type=Instruct2026.02 | 0.8953 | |
| Claude 3 Haiku2026.02 | 0.877 | |
| Qwen1.5 (1.8B)Size=1.8B2026.02 | 0.7488 | |
| Gemma (2B)Size=2B2026.02 | 0.3836 |