Expert-Level Reasoning on XBench-DeepSearch 1.0 (test)
0.9Inference AccuracyReThinker
Evaluation Results
| Method | Links | |
|---|---|---|
| ReThinkerModel Category=Inference Frameworks, Backbone=Gemini-3-Pro, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.9 | |
| Gemini-3-ProModel Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.87 | |
| ReThinkerModel Category=Inference Frameworks, Backbone=OpenPangu-72B, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.78 | |
| GPT-5-highModel Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.778 | |
| Tongyi DeepResearchModel Category=Inference Frameworks, Backbone=30B-A3B, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.75 | |
| DeepSeek-V3.2Model Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.71 | |
| MiroThinker-v1.0Model Category=Inference Frameworks, Backbone=30B, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.706 | |
| GLM-4.6Model Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.7 | |
| Kimi ResearcherModel Category=Inference Frameworks, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.69 | |
| Claude-4.5-SonnetModel Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.66 | |
| WebExplorerModel Category=Inference Frameworks, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.537 | |
| Kimi K2Model Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 0.5 |