Reasoning on TP-Bench
100Accuracy (Level 1)Claude-Opus-4.5
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Claude-Opus-4.5Model Category=Proprietary Models2026.04 | 100 | 98.5 | 76.4 | 41.4 | 9.1 | 63.2 | |
| Gemini-2.5-flashModel Category=Proprietary Models2026.04 | 92.5 | 96.9 | 72.7 | 40.6 | 23.6 | 63.7 | |
| Qwen3-4B-Thinking-2507Model Category=Open-Weight Models2026.04 | 77.5 | 95.4 | 20 | 10 | 0 | 38.9 | |
| OSS-20BModel Category=Open-Weight Models2026.04 | 77.5 | 83.1 | 50.9 | 30 | 1.8 | 47.4 | |
| Qwen3-4B-Instruct-2507Model Category=Open-Weight Models2026.04 | 70 | 78.5 | 3.6 | 7.1 | 0 | 30.2 | |
| DeepSeek-7BModel Category=Open-Weight Models2026.04 | 67.5 | 58.5 | 3.6 | 0 | 0 | 23.5 |