Multi-hop QA on MuSiQue
77.2EMSLEA-RL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SLEA-RLBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 77.2 | — | |
| GSPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 67.6 | — | |
| GRPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 65.1 | — | |
| PPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 63.4 | — | |
| RLOOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 62.2 | — | |
| Reinforce++Base Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 61.2 | — | |
| CoopRAG (Gemma2-27B)Generator Backbone=GPT-4o-mini2025.12 | 52.8 | 66.7 | |
| CoopRAG (Llama3.3-70B)Generator Backbone=GPT-4o-mini2025.12 | 52.6 | 66.6 | |
| CoopRAG (GPT-4o-mini)Generator Backbone=GPT-4o-mini2025.12 | 52.3 | 67.1 | |
| CoopRAG (Gemma2-9B)Generator Backbone=GPT-4o-mini2025.12 | 52.2 | 65.2 | |
| SDGA-PhaseFramework=ZeroSearch, Model=Qwen2.5-7B, Search Depth (Smax)=52026.05 | 51.6 | — | |
| SDGA-AutoFramework=ZeroSearch, Model=Qwen2.5-7B, Search Depth (Smax)=52026.05 | 45.4 | — | |
| SDGA-PhaseFramework=ZeroSearch, Model=Qwen2.5-3B, Search Depth (Smax)=52026.05 | 44.9 | — | |
| HopRAGGenerator Backbone=GPT-4o-mini2025.12 | 42.2 | 54.9 | |
| SiReRAGGenerator Backbone=GPT-4o-mini2025.12 | 40.5 | 53.1 | |
| GRPO-FullFramework=ZeroSearch, Model=Qwen2.5-7B2026.05 | 39.6 | — | |
| SDGA-AutoFramework=ZeroSearch, Model=Qwen2.5-3B, Search Depth (Smax)=52026.05 | 39.4 | — | |
| HippoRAG2Generator Backbone=Llama3.3-70B2025.12 | 37.2 | 48.6 | |
| HippoRAG2Generator Backbone=GPT-4o-mini2025.12 | 35 | 49.3 | |
| NV-Embed-v2 (7B)Generator Backbone=Llama3.3-70B2025.12 | 34.7 | 45.7 | |
| HippoRAG2Backbone=Qwen3-30B-A3B-Instruct-25072026.02 | 34.3 | 45.8 | |
| ToPG-Localmax-iter=32026.01 | 34 | 47 | |
| HELPBackbone=Qwen3-30B-A3B-Instruct-25072026.02 | 33.7 | 44.7 | |
| GritLM-7BGenerator Backbone=Llama3.3-70B2025.12 | 33.6 | 44.8 | |
| GRPO-FullFramework=ZeroSearch, Model=Qwen2.5-3B2026.05 | 32.7 | — | |
| IGPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 31.4 | — | |
| GTE-Qwen2-7B-InstructGenerator Backbone=Llama3.3-70B2025.12 | 30.6 | 40.9 | |
| SDGA-PhaseFramework=Search-R1, Model=Qwen2.5-7B, Search Depth (Smax)=52026.05 | 28.6 | — | |
| GRPO-HalfFramework=ZeroSearch, Model=Qwen2.5-7B2026.05 | 28.2 | — | |
| ToPG-Localmax-iter=12026.01 | 28 | 41.1 | |
| RAPTORGenerator Backbone=GPT-4o-mini2025.12 | 27.7 | 39.2 | |
| GraphRAGGenerator Backbone=Llama3.3-70B2025.12 | 27.3 | 38.5 | |
| GraphRAGGenerator Backbone=GPT-4o-mini2025.12 | 27 | 42 | |
| HippoRAGGenerator Backbone=Llama3.3-70B2025.12 | 26.2 | 35.1 | |
| GTR (T5-base)Generator Backbone=Llama3.3-70B2025.12 | 25.8 | 34.6 | |
| HippoRAG 22026.01 | 24.7 | 36.2 | |
| IRCoT + IndexRAGTime(s)=1.08, Calls=3.22026.03 | 24.6 | — | |
| ContrieverGenerator Backbone=Llama3.3-70B2025.12 | 24 | 31.3 | |
| HippoRAGGenerator Backbone=GPT-4o-mini2025.12 | 24 | 35.9 | |
| HippoRAGTime(s)=3.13, Calls=2.02026.03 | 23.8 | — | |
| SDGA-PhaseFramework=Search-R1, Model=Qwen2.5-3B, Search Depth (Smax)=52026.05 | 23.4 | — | |
| HyperGraphRAGBackbone=Qwen3-30B-A3B-Instruct-25072026.02 | 22.7 | 36.9 | |
| IndexRAGTime(s)=0.30, Calls=1.02026.03 | 22.4 | — | |
| GRPO-HalfFramework=ZeroSearch, Model=Qwen2.5-3B2026.05 | 22.3 | — | |
| SDGA-PhaseFramework=ZeroSearch, Model=LLaMA-3.2-3B-Base, Search Depth (Smax)=52026.05 | 21.4 | — | |
| SDGA-PhaseFramework=Search-R1, Model=LLaMA-3.2-3B-Base, Search Depth (Smax)=52026.05 | 21.1 | — | |
| SDGA-AutoFramework=Search-R1, Model=LLaMA-3.2-3B-Base, Search Depth (Smax)=52026.05 | 20.8 | — | |
| SDGA-AutoFramework=Search-R1, Model=Qwen2.5-7B, Search Depth (Smax)=52026.05 | 20.8 | — | |
| RAPTORGenerator Backbone=Llama3.3-70B2025.12 | 20.7 | 28.9 | |
| BM25Generator Backbone=Llama3.3-70B2025.12 | 20.3 | 28.8 | |
| SkillRLBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 20.2 | — | |
| SDGA-AutoFramework=ZeroSearch, Model=LLaMA-3.2-3B-Base, Search Depth (Smax)=52026.05 | 20.1 | — | |
| Vanilla-RAG2026.01 | 19.7 | 30.6 | |
| ToPG-Naive2026.01 | 19.5 | 30.3 | |
| RAPTORTime(s)=0.47, Calls=1.02026.03 | 19.3 | — | |
| Naive RAGTime(s)=0.29, Calls=1.02026.03 | 19 | — | |
| GiGPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 18.9 | — | |
| LinearRAGBackbone=Qwen3-30B-A3B-Instruct-25072026.02 | 18.7 | 28.1 | |
| ZeroSearchBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 18.4 | — | |
| GraphRAGmode=local2026.01 | 17.8 | 26.7 | |
| FastGraphRAGTime(s)=2.55, Calls=1.02026.03 | 17.5 | — | |
| SDGA-AutoFramework=Search-R1, Model=Qwen2.5-3B, Search Depth (Smax)=52026.05 | 17.1 | — | |
| LightRAGmode=local2026.01 | 16.7 | 25.6 | |
| Search-o1Reasoning Protocol=Retrieval-augmented Reasoning with Reason-in-Documents, Parameter Count=32B2025.01 | 16.6 | 28.2 | |
| GRPO-FullFramework=Search-R1, Model=Qwen2.5-7B2026.05 | 16.4 | — | |
| GRPO-FullFramework=ZeroSearch, Model=LLaMA-3.2-3B-Base2026.05 | 15.8 | — | |
| EvolveRBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 15.6 | — | |
| Llama3.3-70BReasoning Protocol=Direct Reasoning (w/o Retrieval), Parameter Count=70B2025.01 | 14.8 | 23.6 | |
| Search-R1Base Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 14.6 | — | |
| RAgent-QwQ-32BReasoning Protocol=Retrieval-augmented Reasoning, Parameter Count=32B2025.01 | 13.6 | 25.5 | |
| GRPO-FullFramework=Search-R1, Model=LLaMA-3.2-3B-Base2026.05 | 13.6 | — | |
| GRPO-HalfFramework=Search-R1, Model=Qwen2.5-3B2026.05 | 13.4 | — | |
| GRPO-FullFramework=Search-R1, Model=Qwen2.5-3B2026.05 | 13.3 | — | |
| RAgent-Qwen2.5-32BReasoning Protocol=Retrieval-augmented Reasoning, Parameter Count=32B2025.01 | 13 | 25.4 | |
| Qwen2.5-72BReasoning Protocol=Direct Reasoning (w/o Retrieval), Parameter Count=72B2025.01 | 11.4 | 20.4 | |
| SDGA-AntiFramework=Search-R1, Model=Qwen2.5-3B, Search Depth (Smax)=52026.05 | 11.1 | — | |
| RAG-QwQ-32BReasoning Protocol=Retrieval-augmented Reasoning, Parameter Count=32B2025.01 | 10.6 | 20.2 | |
| GRPO-HalfFramework=ZeroSearch, Model=LLaMA-3.2-3B-Base2026.05 | 10.5 | — | |
| RAG-Qwen2.5-32BReasoning Protocol=Retrieval-augmented Reasoning, Parameter Count=32B2025.01 | 10.4 | 19.8 | |
| GRPO-HalfFramework=Search-R1, Model=Qwen2.5-7B2026.05 | 10.2 | — | |
| GRPO-HalfFramework=Search-R1, Model=LLaMA-3.2-3B-Base2026.05 | 10.1 | — | |
| SDGA-AntiFramework=ZeroSearch, Model=LLaMA-3.2-3B-Base, Search Depth (Smax)=52026.05 | 9.8 | — | |
| RAGBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 9.4 | — | |
| SDGA-AntiFramework=Search-R1, Model=LLaMA-3.2-3B-Base, Search Depth (Smax)=52026.05 | 9.2 | — | |
| QwQ-32BReasoning Protocol=Direct Reasoning (w/o Retrieval), Parameter Count=32B2025.01 | 9 | 18.9 | |
| Search-o1Base Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 8.6 | — | |
| Qwen2.5-32BReasoning Protocol=Direct Reasoning (w/o Retrieval), Parameter Count=32B2025.01 | 8.4 | 18 | |
| SDGA-AntiFramework=Search-R1, Model=Qwen2.5-7B, Search Depth (Smax)=52026.05 | 8.4 | — | |
| CoTBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 6.6 | — | |
| R1-InstructBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 6 | — | |
| SDGA-AntiFramework=ZeroSearch, Model=Qwen2.5-7B, Search Depth (Smax)=52026.05 | 5.8 | — | |
| Qwen2.5Base Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 4.8 | — | |
| SDGA-AntiFramework=ZeroSearch, Model=Qwen2.5-3B, Search Depth (Smax)=52026.05 | 4.3 | — | |
| LightRAGGenerator Backbone=GPT-4o-mini2025.12 | 2 | 9.3 | |
| LightRAGGenerator Backbone=Llama3.3-70B2025.12 | 0.5 | 1.6 | |
| IRCoT2026.05 | — | 46.9 | |
| Search-R12026.05 | — | 48.4 | |
| Static RAG2026.05 | — | 42.7 | |
| Step-Levelretrieval_strategy=step-level adaptive retrieval2026.05 | — | 51.2 |