Multi-hop Question Answering on Bamboogle (dev test)
68.2F1 ScoreHiExp-Searcher
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HiExp-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 68.2 | 57.2 | 54.8 | |
| Search-R1-v0.3Evaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 65.1 | 55.2 | 54.4 | |
| GPT-4.1Evaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 63.8 | 55.2 | 49.6 | |
| ReSearchEvaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 63.1 | 52 | 50.4 | |
| DeepSeek-R1Evaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 63 | 52.8 | 52 | |
| o4-miniEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 61.7 | 64 | 46.4 | |
| HiExp-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 61 | 50.4 | 46.4 | |
| Gemini-2.5-ProEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 59.7 | 69.6 | 52 | |
| Search-R1-v0.3Evaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 59.4 | 48 | 47.2 | |
| Qwen3-235B-A22BEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 55.3 | 49.2 | 43.8 | |
| ReSearchEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 53.1 | 45.6 | 41.6 | |
| R1-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 49.8 | 46.4 | 36 | |
| Search-o1 + HiExpEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 44.6 | 37.6 | 33.6 | |
| Search-o1Evaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B, Status=Reproduced2026.04 | 37.5 | 31.2 | 27.2 | |
| IRCoTEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 32.3 | 28.8 | 23.2 | |
| Iter-RetGenEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 31.8 | 24.8 | 22.4 | |
| Vanilla RAGEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 17.6 | 12.8 | 10.4 |