Lawyer Evaluation on OAB-Bench
9.05ScoreGemini-3 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3 ProModel variant=low2026.03 | 9.05 | |
| Gemini-3 ProModel variant=high2026.03 | 8.9 | |
| gpt-5.2Model variant=high2026.03 | 8.73 | |
| gpt-5.2Model variant=instant2026.03 | 8.07 | |
| sabia-42026.03 | 7.49 | |
| gpt-4.12026.03 | 7.3 | |
| sabia-3.12026.03 | 7.21 | |
| sabiazinho-4Price Range=cost-effective2026.03 | 7.02 | |
| kimi-k2Model variant=thinking2026.03 | 6.62 | |
| deepseekModel variant=v3.22026.03 | 6.4 | |
| gpt-5-miniPrice Range=cost-effective2026.03 | 6.37 | |
| Qwen3Model variant=235b2026.03 | 6.33 | |
| gemini-2.5-flash-litePrice Range=cost-effective2026.03 | 6.25 | |
| gpt-oss-120bPrice Range=cost-effective2026.03 | 6.01 | |
| gpt-4.1-miniPrice Range=cost-effective2026.03 | 5.5 |