Search Agent Evaluation on XBench
78Average ScoreDeepSeek-V3.2
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-V3.2Model Category=Foundation Model2026.05 | 78 | |
| GPT-5 HighModel Category=Foundation Model2026.05 | 77 | |
| MiniMax-M2.1Model Category=Foundation Model2026.05 | 68 | |
| Qwen3-8B + ACTGUIDE-RLTraining=ACTGUIDE-RL2026.05 | 44 | |
| Qwen3-4B-Instruct + ACTGUIDE-RLTraining=ACTGUIDE-RL2026.05 | 37 | |
| WebSailor-7BModel Category=Search-Agent-Trained Model2026.05 | 34 | |
| Qwen3-8B + RLTraining=RL Baseline2026.05 | 33 | |
| Qwen3-8BTraining=Base2026.05 | 32 | |
| ARPO-8BModel Category=Search-Agent-Trained Model2026.05 | 25 | |
| WebThinker-32B-RLModel Category=Search-Agent-Trained Model2026.05 | 24 | |
| Qwen2.5-7B-Instruct + ACTGUIDE-RLTraining=ACTGUIDE-RL2026.05 | 24 | |
| Qwen2.5-7B-Instruct + RLTraining=RL Baseline2026.05 | 22 | |
| Qwen2.5-7B-InstructTraining=Base2026.05 | 19 | |
| Qwen3-4B-Instruct + RLTraining=RL Baseline2026.05 | 18 | |
| Qwen2.5-3B-Instruct + ACTGUIDE-RLTraining=ACTGUIDE-RL2026.05 | 16 | |
| Qwen3-4B-InstructTraining=Base2026.05 | 14 | |
| Qwen2.5-3B-Instruct + RLTraining=RL Baseline2026.05 | 10 | |
| Qwen2.5-3B-InstructTraining=Base2026.05 | 8 |