Question Answering on 2WikiMultihopQA (test)
78.9F1KeyExtract + CoopRAG
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| KeyExtract + CoopRAGReader (LLM)=Gemma2-9B2025.12 | 78.9 | 71.5 | — | — | — | — | — | — | — | — | — | — | |
| KeyExtract + CoopRAGReader (LLM)=GPT-4o-mini2025.12 | 78.2 | 72.2 | — | — | — | — | — | — | — | — | — | — | |
| IRCoT + CoopRAGReader (LLM)=GPT-4o-mini2025.12 | 77.7 | 72.2 | — | — | — | — | — | — | — | — | — | — | |
| Graph-R1 + EKABackbone=Qwen2.5-14B-Instruct, Knowledge Interaction=Training, Knowledge Type=Graph-based knowledge2025.12 | 77.12 | 70.31 | — | — | — | — | — | — | — | — | — | — | |
| Graph-R1Backbone=Qwen2.5-14B-Instruct, Knowledge Interaction=Training, Knowledge Type=Graph-based knowledge2025.12 | 75.46 | 67.97 | — | — | — | — | — | — | — | — | — | — | |
| QuCo-RAGLLM Backbone=GPT-4.12025.12 | 74.8 | 64.6 | — | — | — | — | — | — | — | — | — | — | |
| FS-RAGLLM Backbone=GPT-4.12025.12 | 73.8 | 59.5 | — | — | — | — | — | — | — | — | — | — | |
| QuCo-RAGLLM Backbone=GPT-5-chat2025.12 | 73.3 | 59.7 | — | — | — | — | — | — | — | — | — | — | |
| SR-RAGLLM Backbone=GPT-4.12025.12 | 72.6 | 60 | — | — | — | — | — | — | — | — | — | — | |
| IGPOMethod Category=Step-reward RL-based2025.10 | 72.1 | — | — | — | — | — | — | — | — | — | — | — | |
| SR-RAGLLM Backbone=GPT-5-chat2025.12 | 70.1 | 51 | — | — | — | — | — | — | — | — | — | — | |
| Wo-RAGLLM Backbone=GPT-4.12025.12 | 69.9 | 54.7 | — | — | — | — | — | — | — | — | — | — | |
| Web-ToolLLM Backbone=GPT-5-chat2025.12 | 69.8 | 48.3 | — | — | — | — | — | — | — | — | — | — | |
| Graph-R1 + EKABackbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=Graph-based knowledge2025.12 | 68.26 | 60.94 | — | — | — | — | — | — | — | — | — | — | |
| FLARELLM Backbone=GPT-4.12025.12 | 67.9 | 49.8 | — | — | — | — | — | — | — | — | — | — | |
| Wo-RAGLLM Backbone=GPT-5-chat2025.12 | 67 | 50.1 | — | — | — | — | — | — | — | — | — | — | |
| Graph-R1Backbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=Graph-based knowledge2025.12 | 65.04 | 55.47 | — | — | — | — | — | — | — | — | — | — | |
| FS-RAGLLM Backbone=GPT-5-chat2025.12 | 63.3 | 47.3 | — | — | — | — | — | — | — | — | — | — | |
| Web-ToolLLM Backbone=GPT-4.12025.12 | 63.2 | 42.9 | — | — | — | — | — | — | — | — | — | — | |
| IRCoT + HippoRAGReader (LLM)=GPT-3.5-turbo2025.12 | 62.7 | 47.7 | — | — | — | — | — | — | — | — | — | — | |
| Search-R1-PPO + EKABackbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=Chunk-based knowledge2025.12 | 61.47 | 57.03 | — | — | — | — | — | — | — | — | — | — | |
| Search-R1 + EKABackbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=Chunk-based knowledge2025.12 | 60.75 | 56.25 | — | — | — | — | — | — | — | — | — | — | |
| DeepResearcherMethod Category=Outcome-reward RL-based2025.10 | 59.7 | — | — | — | — | — | — | — | — | — | — | — | |
| R1-searcherMethod Category=Outcome-reward RL-based2025.10 | 59.4 | — | — | — | — | — | — | — | — | — | — | — | |
| QuCo-RAGLLM Backbone=Qwen2.5-32B-Instruct2025.12 | 58.9 | 50 | — | — | — | — | — | — | — | — | — | — | |
| ReasoningRAGMethod Category=Step-reward RL-based2025.10 | 50.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Central-GRPOModel=Qwen2.5-3B2026.02 | 50.3 | — | — | — | — | — | — | — | — | 29.3 | — | — | |
| FedGRPOModel=Qwen2.5-3B2026.02 | 49.8 | — | — | — | — | — | — | — | — | 28.5 | — | — | |
| QuCo-RAGLLM Backbone=Llama-3-8B-Instruct2025.12 | 46.6 | 38.4 | — | — | — | — | — | — | — | — | — | — | |
| FS-RAGLLM Backbone=Qwen2.5-32B-Instruct2025.12 | 45.3 | 35.9 | — | — | — | — | — | — | — | — | — | — | |
| IRCoT + ColBERTv2Reader (LLM)=GPT-4o-mini2025.12 | 45.1 | 35.4 | — | — | — | — | — | — | — | — | — | — | |
| StepSearch-baseMethod Category=Step-reward RL-based2025.10 | 45 | — | — | — | — | — | — | — | — | — | — | — | |
| Search-r1-baseMethod Category=Outcome-reward RL-based2025.10 | 44.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GiGPOMethod Category=Step-reward RL-based2025.10 | 43.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Search-r1-instructMethod Category=Outcome-reward RL-based2025.10 | 43.4 | — | — | — | — | — | — | — | — | — | — | — | |
| StepSearch-instructMethod Category=Step-reward RL-based2025.10 | 43.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Search-R1-PPOBackbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=Chunk-based knowledge2025.12 | 42.38 | 39.84 | — | — | — | — | — | — | — | — | — | — | |
| CORE (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=CORE (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 41.62 | 35.86 | — | — | — | — | — | — | — | — | — | 42 | |
| SeaKRLLM Backbone=Llama-3-8B-Instruct2025.12 | 40.4 | 33.5 | — | — | — | — | — | — | — | — | — | — | |
| MHGPO-FoF(os)Framework=LLM + MAS + MARL2025.06 | 40.368 | 32.801 | — | — | — | — | — | — | — | 37.222 | — | — | |
| ETCLLM Backbone=Qwen2.5-32B-Instruct2025.12 | 40.2 | 31.5 | — | — | — | — | — | — | — | — | — | — | |
| CORE (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=CORE (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 39.69 | 33.78 | — | — | — | — | — | — | — | — | — | 34 | |
| MHGPO-RRFramework=LLM + MAS + MARL2025.06 | 39.388 | 31.838 | — | — | — | — | — | — | — | 36.84 | — | — | |
| ETCLLM Backbone=Llama-3-8B-Instruct2025.12 | 39.2 | 29.9 | — | — | — | — | — | — | — | — | — | — | |
| MHGPO-FoFFramework=LLM + MAS + MARL2025.06 | 39.089 | 31.75 | — | — | — | — | — | — | — | 36.299 | — | — | |
| MAPPOFramework=LLM + MAS + MARL2025.06 | 38.722 | 31.767 | — | — | — | — | — | — | — | 35.79 | — | — | |
| Search-R1Backbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=Chunk-based knowledge2025.12 | 38.21 | 35.15 | — | — | — | — | — | — | — | — | — | — | |
| MHGPO-ISFramework=LLM + MAS + MARL2025.06 | 38.05 | 31.115 | — | — | — | — | — | — | — | 34.796 | — | — | |
| Wo-RAGLLM Backbone=Llama-3-8B-Instruct2025.12 | 37.7 | 29.5 | — | — | — | — | — | — | — | — | — | — | |
| R1-Searcher(GRPO)Framework=LLM + RAG (+RL)2025.06 | 37.62 | 30.861 | — | — | — | — | — | — | — | 35.012 | — | — | |
| DRAGINLLM Backbone=Qwen2.5-32B-Instruct2025.12 | 36.9 | 28.8 | — | — | — | — | — | — | — | — | — | — | |
| RECOMP (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=RECOMP (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 36.81 | 30.72 | — | — | — | — | — | — | — | — | — | 44 | |
| FS-RAGLLM Backbone=Llama-3-8B-Instruct2025.12 | 36.8 | 28.8 | — | — | — | — | — | — | — | — | — | — | |
| Top10 DocumentsDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 36.75 | 31.04 | — | — | — | — | — | — | — | — | — | 1,531 | |
| DRAGINLLM Backbone=Llama-3-8B-Instruct2025.12 | 36.7 | 27.9 | — | — | — | — | — | — | — | — | — | — | |
| RECOMP (1B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=RECOMP (1B), Compressor Training Backbone=llama3.2-1B-Instruct2025.08 | 36.53 | 30.45 | — | — | — | — | — | — | — | — | — | 33 | |
| R1-Searcher(PPO)Framework=LLM + RAG (+RL)2025.06 | 36.403 | 29.982 | — | — | — | — | — | — | — | 34.544 | — | — | |
| Top5 DocumentsDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 35.21 | 29.64 | — | — | — | — | — | — | — | — | — | 766 | |
| FLARELLM Backbone=Llama-3-8B-Instruct2025.12 | 35.1 | 26.6 | — | — | — | — | — | — | — | — | — | — | |
| Deepseek-V3 (671B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=Deepseek-V3 (671B)2025.08 | 34.64 | 29 | — | — | — | — | — | — | — | — | — | 40 | |
| R1-SearcherBackbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=Chunk-based knowledge2025.12 | 33.96 | 27.34 | — | — | — | — | — | — | — | — | — | — | |
| FLAREpipeline=FlashRAG2025.05 | 33.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Wo-RAGLLM Backbone=Qwen2.5-32B-Instruct2025.12 | 33.6 | 26.4 | — | — | — | — | — | — | — | — | — | — | |
| Top3 DocumentsDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 33.58 | 27.89 | — | — | — | — | — | — | — | — | — | 460 | |
| FLARELLM Backbone=Qwen2.5-32B-Instruct2025.12 | 33.3 | 26.4 | — | — | — | — | — | — | — | — | — | — | |
| DPSDA-FL+SFTModel=Qwen2.5-3B2026.02 | 32.5 | — | — | — | — | — | — | — | — | 27.2 | — | — | |
| IRCoTpipeline=FlashRAG2025.05 | 32.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-72BFramework=LLM2025.06 | 32.038 | 27.234 | — | — | — | — | — | — | — | 29.373 | — | — | |
| Top1 DocumentDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=RAG without compression2025.08 | 31.87 | 26.79 | — | — | — | — | — | — | — | — | — | 153 | |
| SR-RAGLLM Backbone=Qwen2.5-32B-Instruct2025.12 | 31.8 | 23 | — | — | — | — | — | — | — | — | — | — | |
| SeaKRLLM Backbone=Qwen2.5-32B-Instruct2025.12 | 31.3 | 22.4 | — | — | — | — | — | — | — | — | — | — | |
| R1Backbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=No knowledge interaction2025.12 | 30.99 | 25 | — | — | — | — | — | — | — | — | — | — | |
| Search-o1Method Category=Prompt-based2025.10 | 30.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Deepseek-V3 (671B)Downstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=Deepseek-V3 (671B)2025.08 | 30.31 | 25.07 | — | — | — | — | — | — | — | — | — | 45 | |
| llama3.2-1BDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 10 docs, Compressor Model=llama3.2-1B2025.08 | 30.06 | 24.93 | — | — | — | — | — | — | — | — | — | 61 | |
| llama3.2-1BDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=Compression of top 5 docs, Compressor Model=llama3.2-1B2025.08 | 30.03 | 24.98 | — | — | — | — | — | — | — | — | — | 61 | |
| No RetrievalDownstream LLM (M)=Qwen2.5-14B-Instruct, Retrieval/Compression Setting=No Retrieval2025.08 | 29.51 | 26.11 | — | — | — | — | — | — | — | — | — | 0 | |
| SR-RAGLLM Backbone=Llama-3-8B-Instruct2025.12 | 29.2 | 12.9 | — | — | — | — | — | — | — | — | — | — | |
| DioRLLM=Qwen2.5-7B, Retrieval=BM252025.04 | 29.02 | 21.4 | 29.13 | 30.32 | 2.036 | 2.04 | 1.019 | 442.793 | 28.751 | — | — | — | |
| DPSDA-FL+GRPOModel=Qwen2.5-3B2026.02 | 28.8 | — | — | — | — | — | — | — | — | 25.4 | — | — | |
| RaDIOLLM=Qwen2.5-7B, Retrieval=BM252025.04 | 26.41 | 17.8 | 27.04 | 25.56 | 1.015 | 2.301 | 1.135 | 521.979 | 36.474 | — | — | — | |
| CoTMethod Category=Prompt-based2025.10 | 26.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Fedpetuning+SFTModel=Qwen2.5-3B2026.02 | 25.8 | — | — | — | — | — | — | — | — | 19.2 | — | — | |
| Search-o1Framework=LLM + RAG (+RL)2025.06 | 25.68 | 17.802 | — | — | — | — | — | — | — | 24.231 | — | — | |
| Qwen2.5-72BFramework=LLM + MAS2025.06 | 25.63 | 18.511 | — | — | — | — | — | — | — | 36.379 | — | — | |
| BaseLLM=Qwen2.5-7B, Retrieval=BM252025.04 | 24.6 | 16.5 | 25.9 | 24.5 | 1.866 | 3.737 | 1.866 | 856.514 | 54.754 | — | — | — | |
| CoT+RAGMethod Category=Prompt-based2025.10 | 24.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Central-GRPOModel=Qwen2.5-1.5B2026.02 | 24.1 | — | — | — | — | — | — | — | — | 27.1 | — | — | |
| FedGRPOModel=Qwen2.5-1.5B2026.02 | 23.7 | — | — | — | — | — | — | — | — | 27.6 | — | — | |
| Fedpetuning+GRPOModel=Qwen2.5-3B2026.02 | 23.5 | — | — | — | — | — | — | — | — | 18.5 | — | — | |
| Llama3.1-8BFramework=LLM + MAS2025.06 | 22.427 | 15.872 | — | — | — | — | — | — | — | 22.281 | — | — | |
| StandardRAGBackbone=GPT-4o-mini, Knowledge Interaction=Prompt Engineering, Knowledge Type=Chunk-based knowledge2025.12 | 22.31 | 7.03 | — | — | — | — | — | — | — | — | — | — | |
| HyperGraphRAGBackbone=GPT-4o-mini, Knowledge Interaction=Prompt Engineering, Knowledge Type=Graph-based knowledge2025.12 | 21.14 | 4.69 | — | — | — | — | — | — | — | — | — | — | |
| REPLUGpipeline=FlashRAG2025.05 | 21.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Standard RAGpipeline=FlashRAG2025.05 | 21 | — | — | — | — | — | — | — | — | — | — | — | |
| Llama3.1-8BFramework=LLM2025.06 | 20.817 | 7.069 | — | — | — | — | — | — | — | 16.963 | — | — | |
| SuRepipeline=FlashRAG2025.05 | 20.6 | — | — | — | — | — | — | — | — | — | — | — | |
| SFTBackbone=Qwen2.5-7B-Instruct, Knowledge Interaction=Training, Knowledge Type=No knowledge interaction2025.12 | 20.28 | 11.72 | — | — | — | — | — | — | — | — | — | — | |
| NaiveGenerationBackbone=GPT-4o-mini, Knowledge Interaction=Prompt Engineering, Knowledge Type=No knowledge interaction2025.12 | 17.03 | 4.69 | — | — | — | — | — | — | — | — | — | — | |
| LightRAGBackbone=GPT-4o-mini, Knowledge Interaction=Prompt Engineering, Knowledge Type=Graph-based knowledge2025.12 | 16.59 | 3.13 | — | — | — | — | — | — | — | — | — | — |