Code Agent Simulation on Terminal Bench 2.0
54.2AccuracyGemini-3.0 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3.0 Pro2025.12 | 54.2 | |
| DeepSeek-V3.2thinking mode=true2025.12 | 46.4 | |
| Claude-4.5-Sonnet2025.12 | 42.8 | |
| Kimi-K2thinking mode=true2025.12 | 35.7 | |
| GPT-5 High2025.12 | 35.2 | |
| MiniMax M22025.12 | 30 |