Deep Search Question Answering on Deep SearchQA
80ScoreClaude-4.5-Opus
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude-4.5-OpusModel Category=Foundation Models with Tools2026.03 | 80 | |
| OpenAI GPT-5 HighModel Category=Foundation Models with Tools2026.03 | 79 | |
| Kimi-K2.5Model Category=Foundation Models with Tools2026.03 | 77.1 | |
| Gemini-3.0-ProModel Category=Foundation Models with Tools2026.03 | 76.9 | |
| MiroThinker-v1.7-miniModel Category=Trained Agents (≥30B)2026.03 | 67.9 | |
| DeepSeek-V3.2Model Category=Foundation Models with Tools2026.03 | 60.9 | |
| MiroThinker-v1.0-8BModel Category=Trained Agents (≤8B), Evaluated using our implementation=true2026.03 | 36.7 | |
| AgentCPM-Explore-4BModel Category=Trained Agents (≤8B), Evaluated using our implementation=true2026.03 | 32.8 | |
| Marco-DR-8BModel Category=Trained Agents (≤8B)2026.03 | 29.2 | |
| WebExplorer-8B-RLModel Category=Trained Agents (≤8B), Evaluated using our implementation=true2026.03 | 17.8 |