Multi-hop Question Answering on HotpotQA (dev test)
71.2F1 ScoreHiExp-Searcher
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HiExp-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 71.2 | 62.9 | 57.8 | |
| ReSearchEvaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 69.4 | 61 | 56.3 | |
| Search-R1-v0.3Evaluation Protocol=Training Based, Base Model=Qwen2.5-32B2026.04 | 66.5 | 55.8 | 53.5 | |
| HiExp-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 65.4 | 60.4 | 52.4 | |
| ReSearchEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 63.2 | 55.8 | 50.4 | |
| DeepSeek-R1Evaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 62.5 | 54 | 48 | |
| Search-R1-v0.3Evaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 61.8 | 53.6 | 49.8 | |
| GPT-4.1Evaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 60.6 | 56 | 45 | |
| o4-miniEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 57.8 | 59.5 | 40.5 | |
| R1-SearcherEvaluation Protocol=Training Based, Base Model=Qwen2.5-7B2026.04 | 57.8 | 59.7 | 45.6 | |
| Qwen3-235B-A22BEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 57.3 | 56.1 | 44.5 | |
| Gemini-2.5-ProEvaluation Protocol=Prompt Based, Category=Frontier LLMs2026.04 | 55.6 | 60.5 | 39.5 | |
| Iter-RetGenEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 51.4 | 45.2 | 39.9 | |
| Search-o1 + HiExpEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 48.7 | 47.7 | 37.1 | |
| IRCoTEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 47.2 | 47.3 | 35.3 | |
| Search-o1Evaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B, Status=Reproduced2026.04 | 44.4 | 41.2 | 34.2 | |
| Vanilla RAGEvaluation Protocol=Prompt Based, Base Model=Qwen2.5-7B2026.04 | 29 | 22.4 | 20.5 |