Multiple Choice Question Answering on EXAMS
95AccuracyGemini-3 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3 ProModel variant=high2026.03 | 95 | |
| Gemini-3 ProModel variant=low2026.03 | 93.3 | |
| gpt-5.2Model variant=high2026.03 | 92.9 | |
| gpt-5.2Model variant=instant2026.03 | 88 | |
| sabia-42026.03 | 86.6 | |
| gpt-4.12026.03 | 86.1 | |
| gpt-5-miniPrice Range=cost-effective2026.03 | 84.6 | |
| deepseekModel variant=v3.22026.03 | 84 | |
| kimi-k2Model variant=thinking2026.03 | 83 | |
| sabia-3.12026.03 | 82.4 | |
| Qwen3Model variant=235b2026.03 | 82 | |
| sabiazinho-4Price Range=cost-effective2026.03 | 81 | |
| gpt-4.1-miniPrice Range=cost-effective2026.03 | 81 | |
| gpt-oss-120bPrice Range=cost-effective2026.03 | 77 | |
| gemini-2.5-flash-litePrice Range=cost-effective2026.03 | 76.2 | |
| Gemma 2shot_count=5-shot2026.03 | 71.2 | |
| TildeOpen LLMshot_count=5-shot2026.03 | 66.6 | |
| ALIAshot_count=5-shot2026.03 | 62.7 | |
| EuroLLMshot_count=5-shot2026.03 | 62.5 | |
| GPT-4Setting=Few-shot2024.12 | 57.76 | |
| LLaMA3-Tamed-70BSetting=Few-shot2024.12 | 55.49 | |
| Llama3-70BSetting=Few-shot2024.12 | 54.78 | |
| Qwen1.5-32BSetting=Few-shot2024.12 | 52.01 | |
| Qwen1.5-72BSetting=Few-shot2024.12 | 48.68 | |
| Llama3-8BSetting=Few-shot2024.12 | 46.34 | |
| LLaMA3-Tamed-8BSetting=Few-shot2024.12 | 46.15 | |
| ChatGPT 3.5 TurboSetting=Few-shot2024.12 | 45.93 | |
| Jais-30B-v3Setting=Few-shot2024.12 | 45.78 | |
| Qwen1.5-7BSetting=Few-shot2024.12 | 38.34 |