Multi-hop Question Answering on 2WikiMultiHopQA (dev test)
81.5F1 ScoreHiExp-Searcher
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HiExp-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 81.5 | 80.4 | 75.8 | |
| ReSearchEvaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 78.1 | 76.7 | 72.3 | |
| HiExp-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 74.6 | 76.5 | 66.9 | |
| Search-R1-v0.3Evaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 73.4 | 71.7 | 68.1 | |
| Gemini-2.5-ProEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 71.8 | 83 | 60.5 | |
| GPT-4.1Evaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 69.7 | 75.5 | 56 | |
| ReSearchEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 67.1 | 65.4 | 60.3 | |
| DeepSeek-R1Evaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 65.7 | 65 | 54 | |
| R1-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 64 | 67.8 | 56.2 | |
| o4-miniEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 62.1 | 71 | 47.5 | |
| Search-R1-v0.3Evaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 60.7 | 58.7 | 52.3 | |
| Qwen3-235B-A22BEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 59.4 | 64.1 | 45.3 | |
| Search-o1 + HiExpEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 54.8 | 56.8 | 45.2 | |
| Search-o1Evaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B, Status=Reproduced2026.04 | 50.8 | 51 | 41.8 | |
| Iter-RetGenEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 39.2 | 35.5 | 32.2 | |
| IRCoTEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 35 | 39.2 | 25.5 | |
| Vanilla RAGEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 32.5 | 27.9 | 27 |