Deep Search on xBench-DeepSearch DS-2510
75ScoreOpenAI GPT-5 High
Evaluation Results
| Method | Links | |
|---|---|---|
| OpenAI GPT-5 HighModel Category=Foundation Models with Tools2026.03 | 75 | |
| MiroThinker-v1.7-miniModel Category=Trained Agents (≥30B)2026.03 | 57.2 | |
| DeepSeek-V3.2Model Category=Foundation Models with Tools2026.03 | 55.7 | |
| Tongyi-DR-30BModel Category=Trained Agents (≥30B)2026.03 | 55 | |
| Gemini-3.0-ProModel Category=Foundation Models with Tools2026.03 | 53 | |
| GLM-4.7Model Category=Foundation Models with Tools2026.03 | 52.3 | |
| Kimi-K2.5Model Category=Foundation Models with Tools2026.03 | 46 | |
| Minimax-M2.1Model Category=Foundation Models with Tools2026.03 | 43 | |
| Marco-DR-8BModel Category=Trained Agents (≤8B)2026.03 | 42 | |
| AgentCPM-Explore-4BModel Category=Trained Agents (≤8B), Evaluated using our implementation=true2026.03 | 34 | |
| MiroThinker-v1.0-8BModel Category=Trained Agents (≤8B), Evaluated using our implementation=true2026.03 | 34 | |
| WebExplorer-8B-RLModel Category=Trained Agents (≤8B), Evaluated using our implementation=true2026.03 | 23 |