Deep Research on BrowseComp+
55.33AccuracyQwen3-235B (w/ Pensieve)
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-235B (w/ Pensieve)Context=256K2026.02 | 55.33 | |
| OPENRESEARCHER2026.03 | 54.8 | |
| StateLM-14B-RLContext=128K2026.02 | 52.67 | |
| StateLM-14BContext=128K2026.02 | 51.33 | |
| StateLM-8B-RLContext=128K2026.02 | 46.44 | |
| StateLM-8BContext=128K2026.02 | 46.22 | |
| Tongyi DeepResearchtype=Deep Research Agents2026.03 | 44.5 | |
| Claude-4-Opustype=Foundation Models with Tools2026.03 | 36.8 | |
| GPT-4.1type=Foundation Models with Tools2026.03 | 36.4 | |
| Kimi-K2type=Foundation Models with Tools2026.03 | 35.4 | |
| StateLM-4BContext=128K2026.02 | 35.33 | |
| CutBill-30B-A3Btype=Deep Research Agents2026.03 | 30.3 | |
| Gemini-2.5-ProScaffold=ReAct2026.02 | 29.52 | |
| Gemini-2.5-Protype=Foundation Models with Tools2026.03 | 29.5 | |
| WideSeek-8B-RLScaffold=WideSeek2026.02 | 26.42 | |
| GPT-OSS-120B-LowScaffold=ReAct2026.02 | 25.54 | |
| WideSeek-8B-SFTScaffold=WideSeek2026.02 | 23.61 | |
| WideSeek-8B-SFT-RLScaffold=WideSeek2026.02 | 23.61 | |
| Nemotron-3-Nanotype=Foundation Models with Tools2026.03 | 20.8 | |
| DeepSeek-R1type=Foundation Models with Tools2026.03 | 16.4 | |
| DeepSeek-R1-0528Scaffold=ReAct2026.02 | 16.39 | |
| Qwen3-30B-A3BScaffold=WideSeek2026.02 | 14.82 | |
| Qwen3-8BScaffold=WideSeek2026.02 | 14.22 | |
| Search-R1-32BScaffold=ReAct2026.02 | 11.08 | |
| Qwen3-32BScaffold=ReAct2026.02 | 10.72 | |
| Qwen3-4B-Instruct (TIPS)Model=Qwen3-4B-Instruct, Training Protocol=TIPS2026.03 | 9.4 | |
| Qwen3-4B-Instruct (PPO)Model=Qwen3-4B-Instruct, Training Protocol=PPO2026.03 | 6.75 | |
| Qwen3-8BContext=128K2026.02 | 5.56 | |
| Qwen3-14BContext=128K2026.02 | 5.46 | |
| Search-R1-32BModel=Search-R1-32B, Training Protocol=PPO2026.03 | 4.11 | |
| Qwen2.5-7B-Instruct (TIPS)Model=Qwen2.5-7B-Instruct, Training Protocol=TIPS2026.03 | 4.1 | |
| Qwen2.5-7B-Instruct (GRPO)Model=Qwen2.5-7B-Instruct, Training Protocol=GRPO2026.03 | 3.61 | |
| Qwen3-4BContext=128K2026.02 | 2.89 | |
| Qwen3-4B-Instruct (GRPO)Model=Qwen3-4B-Instruct, Training Protocol=GRPO2026.03 | 2.65 | |
| Qwen2.5-3B-Instruct (TIPS)Model=Qwen2.5-3B-Instruct, Training Protocol=TIPS2026.03 | 1.57 | |
| Qwen2.5-7B-Instruct (PPO)Model=Qwen2.5-7B-Instruct, Training Protocol=PPO2026.03 | 1.24 | |
| Qwen2.5-3B-Instruct (PPO)Model=Qwen2.5-3B-Instruct, Training Protocol=PPO2026.03 | 1.2 | |
| Qwen2.5-3B-Instruct (GRPO)Model=Qwen2.5-3B-Instruct, Training Protocol=GRPO2026.03 | 1.2 |