Judge Evaluation on Magis-Bench
7.79ScoreGemini-3 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3 ProModel variant=low2026.03 | 7.79 | |
| Gemini-3 ProModel variant=high2026.03 | 7.48 | |
| gpt-5.2Model variant=high2026.03 | 6.99 | |
| gpt-5.2Model variant=instant2026.03 | 6.66 | |
| gpt-4.12026.03 | 5.55 | |
| sabia-42026.03 | 5.08 | |
| sabia-3.12026.03 | 4.97 | |
| deepseekModel variant=v3.22026.03 | 4.88 | |
| Qwen3Model variant=235b2026.03 | 4.52 | |
| sabiazinho-4Price Range=cost-effective2026.03 | 4.5 | |
| kimi-k2Model variant=thinking2026.03 | 4.49 | |
| gpt-5-miniPrice Range=cost-effective2026.03 | 4.47 | |
| gemini-2.5-flash-litePrice Range=cost-effective2026.03 | 4.25 | |
| gpt-4.1-miniPrice Range=cost-effective2026.03 | 3.67 | |
| gpt-oss-120bPrice Range=cost-effective2026.03 | 3.62 |