Deep Search Reasoning on XBench DeepSearch2505
41ScoreClaude-3.7-Sonnet
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude-3.7-SonnetModel Type=Proprietary2026.02 | 41 | |
| CSOBase Model=CK-Pro-8B, Post-Training Method=Critical Step Optimization2026.02 | 29 | |
| GPT-4.1Model Type=Proprietary2026.02 | 27 | |
| Step-DPOBase Model=CK-Pro-8B, Post-Training Method=Step-wise DPO2026.02 | 25 | |
| IPRBase Model=CK-Pro-8B, Post-Training Method=Iterative Process Refinement2026.02 | 24 | |
| CK-Pro-8BMode=SFT2026.02 | 23 | |
| ETOBase Model=CK-Pro-8B, Post-Training Method=Exploration-based Trajectory Optimization2026.02 | 22 | |
| RFTBase Model=CK-Pro-8B, Post-Training Method=Rejection Sampling Fine-Tuning2026.02 | 20 | |
| Qwen3-8BModel Type=Open-Source Baseline2026.02 | 7 |