Deep Research on xBench DS 2510
75ScoreGPT-5 High
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5 HighModel Category=Foundation Models2026.04 | 75 | |
| DeepSeek-V3.2Model Category=Foundation Models2026.04 | 55.7 | |
| Tongyi-DR-30BModel Category=Trained Agents, Model Scale Group=≥30B2026.04 | 55 | |
| Gemini-3-ProModel Category=Foundation Models2026.04 | 53 | |
| GLM-4.7Model Category=Foundation Models2026.04 | 52.3 | |
| Kimi-K2.5Model Category=Foundation Models2026.04 | 46 | |
| MiniMax-M2.1Model Category=Foundation Models2026.04 | 43 | |
| DR-Venus-4B-RLModel Category=Trained Agents, Model Scale Group=≤9B, Training Protocol=RL2026.04 | 40.7 | |
| DR-Venus-4B-SFTModel Category=Trained Agents, Model Scale Group=≤9B, Training Protocol=SFT2026.04 | 35.3 | |
| AgentCPM-Explore-4BModel Category=Trained Agents, Model Scale Group=≤9B2026.04 | 34 | |
| WebExplorer-8B-RLModel Category=Trained Agents, Model Scale Group=≤9B, Training Protocol=RL2026.04 | 23 |