Multi-Hop Question Answering on HotpotQA in-domain
46.3EMSIGHT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SIGHTModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 46.3 | 1.73 | |
| ARPOModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 46.1 | 2.18 | |
| EM+Info-Think+Think-AnsFramework=VERITAS-R1, Reward Components=EM+Info-Think+Think-Ans2025.10 | 44.6 | — | |
| EM+Info-ThinkFramework=VERITAS-R1, Reward Components=EM+Info-Think2025.10 | 44.5 | — | |
| EM+Think-AnsFramework=VERITAS-R1, Reward Components=EM+Think-Ans2025.10 | 44.5 | — | |
| RECONBackbone=Qwen2.5-7B-Base, Training Protocol=PPO, RECON Stage=Stage 1 + Stage 22025.10 | 44.5 | — | |
| Search-R1-7B-Base-PPO w/ FormatModel Scale=7B, Training Algorithm=PPO, Input Formatting=Included2025.10 | 43.6 | — | |
| ReSearch-7B-InstructModel Scale=7B, Model Variant=Instruct2025.10 | 43.5 | — | |
| Search-R1-7B-Base-PPOModel Scale=7B, Training Algorithm=PPO2025.10 | 43.3 | — | |
| Search-R1Backbone=Qwen2.5-7B-Base, Training Protocol=PPO2025.10 | 43.3 | — | |
| DeSA-7B-Instrct-GRPOModel Scale=7B, Training Algorithm=GRPO2025.10 | 42.4 | — | |
| SIGHTModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 41.99 | 1.83 | |
| ARPOModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 41.6 | 3.06 | |
| ReSearch-7B-BaseModel Scale=7B, Model Variant=Base2025.10 | 40.6 | — | |
| Tree-GRPOModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 40.43 | 1.7 | |
| Search-R1Model=Qwen2.5-7B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 36.91 | 2.95 | |
| Search-R1Model=Qwen2.5-3B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 35.74 | 2.31 | |
| Search-o1Model=Qwen2.5-7B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 35.7 | 1.39 | |
| DeSA-3B-Instruct-GRPOModel Scale=3B, Training Algorithm=GRPO2025.10 | 35.2 | — | |
| Rejection SamplingReasoning Strategy=Rejection Sampling2025.10 | 33.1 | — | |
| Rejection Sampling2025.10 | 33.1 | — | |
| Tree-GRPOModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 32.62 | 1.76 | |
| RECONBackbone=Qwen2.5-3B-Base, Training Protocol=PPO, RECON Stage=Stage 1 + Stage 22025.10 | 32.6 | — | |
| RECONBackbone=Qwen2.5-7B-Base, Training Protocol=PPO, RECON Stage=Stage 2 Distillation-only2025.10 | 30.5 | — | |
| RAGReasoning Strategy=Retrieval-Augmented Generation2025.10 | 29.9 | — | |
| RAG2025.10 | 29.9 | — | |
| Search-R1Backbone=Qwen2.5-3B-Base, Training Protocol=PPO2025.10 | 28.4 | — | |
| RECONBackbone=Qwen2.5-3B-Base, Training Protocol=PPO, RECON Stage=Stage 2 Distillation-only2025.10 | 27.1 | — | |
| Search-o1Model=Qwen2.5-3B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 25.2 | 1.19 | |
| R1-baseModel Variant=Base2025.10 | 24.2 | — | |
| R1-base2025.10 | 24.2 | — | |
| GRPOModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/o Retrieval2026.02 | 23.83 | 0 | |
| R1-instructModel Variant=Instruct2025.10 | 23.7 | — | |
| R1-instruct2025.10 | 23.7 | — | |
| SFTTraining Algorithm=Supervised Fine-Tuning2025.10 | 21.7 | — | |
| SFT2025.10 | 21.7 | — | |
| ReActModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 18.95 | 1.91 | |
| Search-o1Reasoning Strategy=Search-o12025.10 | 18.7 | — | |
| Search-o12025.10 | 18.7 | — | |
| Direct InferenceReasoning Strategy=Direct Inference2025.10 | 18.3 | — | |
| Direct Inference2025.10 | 18.3 | — | |
| GRPOModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/o Retrieval2026.02 | 17.38 | 0 | |
| Vanilla RAGModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 14.8 | 1 | |
| IRCoTReasoning Strategy=Iterative Retrieval Chain-of-Thought2025.10 | 13.3 | — | |
| IRCoT2025.10 | 13.3 | — | |
| CoTReasoning Strategy=Chain-of-Thought2025.10 | 9.2 | — | |
| CoT2025.10 | 9.2 | — | |
| Direct ReasoningModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/o Retrieval2026.02 | 7.42 | 0 | |
| Vanilla RAGModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 5.66 | 1 | |
| COTModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/o Retrieval2026.02 | 4.1 | 0 | |
| ReActModel=Qwen2.5-3B-Instruct, Retrieval Configuration=w/ Retrieval2026.02 | 2.54 | 3.12 | |
| Direct ReasoningModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/o Retrieval2026.02 | 0 | 0 | |
| COTModel=Qwen2.5-7B-Instruct, Retrieval Configuration=w/o Retrieval2026.02 | 0 | 0 |