Question Answering on PopQA (EM, F1)
56.6EMT^2RAG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| T^2RAGInference LLM=Gemini-2.5-flash2025.08 | 56.6 | 62.4 | |
| T^2RAGInference LLM=GPT-4o-mini2025.08 | 55.8 | 63.2 | |
| RAPTORInference LLM=GPT-4o-mini2025.08 | 54.6 | 60.1 | |
| RAPTORInference LLM=Gemini-2.5-flash2025.08 | 52.3 | 56.8 | |
| HippoRAG2Inference LLM=GPT-4o-mini2025.08 | 52.2 | 60.2 | |
| HippoRAG2Inference LLM=Gemini-2.5-flash2025.08 | 52.1 | 60.1 | |
| StandardInference LLM=GPT-4o-mini2025.08 | 51.9 | 60 | |
| StandardInference LLM=Gemini-2.5-flash2025.08 | 51.8 | 59.5 | |
| ToPG-Naive2026.01 | 51.6 | 63.9 | |
| DecEx-RAGprotocol=Learning-Based2026.01 | 51.3 | 53.2 | |
| Search-R1Backbone=Qwen2.5-7B2025.12 | 51.3 | 57.1 | |
| IRCoTInference LLM=Gemini-2.5-flash2025.08 | 51.2 | 58.7 | |
| RouteRAGBackbone=Qwen2.5-7B2025.12 | 50.6 | 56.4 | |
| BM25Inference LLM=Gemini-2.5-flash2025.08 | 50.2 | 55.6 | |
| RouteRAGBackbone=Qwen2.5-3B2025.12 | 49.4 | 56.8 | |
| Vanilla-RAG2026.01 | 49.2 | 62.2 | |
| ToPG-Localmax-iter=32026.01 | 48.9 | 60.2 | |
| BAR-RAGBackbone=LLaMA-3.1-8B-Instruct, Number of Iterations=32026.02 | 48.6 | — | |
| ToPG-Localmax-iter=12026.01 | 48.4 | 59.5 | |
| BAR-RAGBackbone=LLaMA-3.1-8B-Instruct, Number of Iterations=22026.02 | 48 | — | |
| BM25Inference LLM=GPT-4o-mini2025.08 | 47.6 | 54.8 | |
| DPR + IRCoT + Q3RType=Traditional RAG, Reranker=Qwen3-Reranker-8B (Q3R), Iterative Retrieval Protocol=IRCoT2026.05 | 47.3 | 56.98 | |
| ProRAGMethod Category=Reinforcement Learning-based Methods2026.01 | 47.2 | 51.6 | |
| DPRType=Traditional RAG2026.05 | 47.2 | 55.89 | |
| BM25 + DPR + Q3RType=Traditional RAG, Reranker=Qwen3-Reranker-8B (Q3R)2026.05 | 47.2 | 56.01 | |
| Search-o1protocol=Prompt-Based2026.01 | 47 | 50 | |
| BAR-RAGBackbone=Qwen-2.5-7B-Instruct, Number of Iterations=32026.02 | 46.9 | — | |
| Ψ-RAG ⍋Type=Tree-RAG, Reranker=Qwen3-Reranker-8B (Q3R), Abstract Type=summative2026.05 | 46.7 | 56.74 | |
| BAR-RAGBackbone=Qwen-2.5-7B-Instruct, Number of Iterations=22026.02 | 46.3 | — | |
| Ψ-RAG øType=Tree-RAG, Reranker=Qwen3-Reranker-8B (Q3R), Abstract Type=keyword2026.05 | 46.1 | 55.88 | |
| Search-R1Method Category=Reinforcement Learning-based Methods2026.01 | 45.9 | 50.1 | |
| BAR-RAGBackbone=LLaMA-3.1-8B-Instruct, Number of Iterations=12026.02 | 45.8 | — | |
| Search-R1Backbone=Qwen2.5-3B2025.12 | 45.8 | 53.3 | |
| IRCoTInference LLM=GPT-4o-mini2025.08 | 45.3 | 54.7 | |
| BAR-RAGBackbone=Qwen-2.5-3B-Instruct, Number of Iterations=32026.02 | 44.1 | — | |
| BAR-RAGBackbone=Qwen-2.5-7B-Instruct, Number of Iterations=12026.02 | 44.1 | — | |
| BAR-RAGBackbone=Qwen-2.5-3B-Instruct, Number of Iterations=22026.02 | 43.4 | — | |
| HippoRAG 2Type=Graph-RAG2026.05 | 43.4 | 55.9 | |
| HippoRAG 2 + Q3RType=Graph-RAG, Reranker=Qwen3-Reranker-8B (Q3R)2026.05 | 43.3 | 56.05 | |
| Iter-RetGenprotocol=Prompt-Based2026.01 | 42.5 | 49.3 | |
| HippoRAGBackbone=GPT-4o-mini2025.12 | 42.5 | 56.2 | |
| RAPTORBackbone=GPT-4o-mini2025.12 | 41.9 | 55.1 | |
| SuReBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs2026.03 | 41.8 | 48.99 | |
| HippoRAG 2Backbone=GPT-4o-mini2025.12 | 41.7 | 55.7 | |
| ReasonRAGMethod Category=Reinforcement Learning-based Methods2026.01 | 41.5 | 46.2 | |
| RAG SFTBackbone=Qwen-2.5-3B-Instruct2026.02 | 41.4 | — | |
| Search-R1protocol=Learning-Based2026.01 | 41.3 | 46.4 | |
| BAR-RAGBackbone=Qwen-2.5-3B-Instruct, Number of Iterations=12026.02 | 41.2 | — | |
| ReasonRAGprotocol=Learning-Based2026.01 | 41.1 | 44.4 | |
| BM25Type=Traditional RAG2026.05 | 40.9 | 49.01 | |
| DeepRAGprotocol=Learning-Based2026.01 | 40.6 | 43.2 | |
| RECOMPprotocol=Prompt-Based2026.01 | 40.5 | 45.8 | |
| LightRAGmode=local2026.01 | 39.7 | 53.4 | |
| RAG w/RerankerBackbone=Qwen-2.5-3B-Instruct2026.02 | 39.6 | — | |
| Iter-RetGenBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs2026.03 | 39.6 | 46.41 | |
| FLAREMethod Category=Advanced Methods2026.01 | 39.3 | 46.3 | |
| LongLLMLinguaprotocol=Prompt-Based2026.01 | 39.2 | 45.1 | |
| RAGShaperprotocol=Learning-Based, training_data_size=6.5k2026.01 | 38.9 | 49.6 | |
| VOTE-RAGBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs, parallel agents (N)=32026.03 | 38.9 | 46.8 | |
| RAGBackbone=Qwen-2.5-3B-Instruct2026.02 | 38.7 | — | |
| IKEAprotocol=Learning-Based2026.01 | 38.7 | 42.7 | |
| DRAGBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs, max debate interactions (r)=3, agents=32026.03 | 38.6 | 46.5 | |
| Search-o1Method Category=Advanced Methods2026.01 | 38.4 | 43.9 | |
| HippoRAG 22026.01 | 38.4 | 48.6 | |
| GraphRAGmode=local2026.01 | 38.1 | 52.6 | |
| Naive RAGBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs2026.03 | 37.6 | 45.69 | |
| RAGShaperprotocol=Learning-Based, training_data_size=4.5k2026.01 | 37.4 | 47.8 | |
| FLAREprotocol=Prompt-Based2026.01 | 36.8 | 44.9 | |
| RAG SFTBackbone=LLaMA-3.1-8B-Instruct2026.02 | 35.6 | — | |
| HL-Dataprotocol=Learning-Based, training_data_size=4.5k2026.01 | 35.2 | 48.3 | |
| Selective-Contextprotocol=Prompt-Based2026.01 | 34.9 | 41.5 | |
| HiPRAGMethod Category=Reinforcement Learning-based Methods2026.01 | 34.3 | 42.7 | |
| IRCoTMethod Category=Advanced Methods2026.01 | 33.6 | 42.2 | |
| Iter-RetGenMethod Category=Advanced Methods2026.01 | 32.9 | 37.8 | |
| GraphRAGType=Graph-RAG2026.05 | 32.7 | 42.13 | |
| IR-COTprotocol=Prompt-Based2026.01 | 32.4 | 39.9 | |
| NORInference LLM=Gemini-2.5-flash2025.08 | 32.4 | 35.7 | |
| RAG SFTBackbone=Qwen-2.5-7B-Instruct2026.02 | 32.3 | — | |
| PixSearchParameters=13B, Search Mode=Full2026.01 | 31.6 | 32.18 | |
| IRCoTBackbone=LLaMA-3.1-8B-Instruct2026.02 | 31.2 | — | |
| GraphRAGBackbone=GPT-4o-mini2025.12 | 30.7 | 51.3 | |
| RAG w/RerankerBackbone=LLaMA-3.1-8B-Instruct2026.02 | 30.4 | — | |
| Vanilla RAGBackbone=Qwen2.5-3B2025.12 | 30.3 | 41.6 | |
| IRCoTBackbone=Qwen-2.5-7B-Instruct2026.02 | 30.1 | — | |
| HippoRAG 2Backbone=Qwen2.5-3B2025.12 | 29.1 | 40.1 | |
| RAGBackbone=LLaMA-3.1-8B-Instruct2026.02 | 28.9 | — | |
| NORInference LLM=GPT-4o-mini2025.08 | 28.7 | 31.4 | |
| R1-SearcherBackbone=Qwen2.5-7B2025.12 | 28.4 | 41 | |
| RAG w/RerankerBackbone=Qwen-2.5-7B-Instruct2026.02 | 27.3 | — | |
| Standard RAGMethod Category=Standard Methods2026.01 | 27 | 35.9 | |
| HippoRAG 2Backbone=Qwen2.5-7B2025.12 | 27 | 37.9 | |
| IRCoTBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs2026.03 | 27 | 33.02 | |
| RAGBackbone=Qwen-2.5-7B-Instruct2026.02 | 26.7 | — | |
| Vanilla RAGBackbone=Qwen2.5-7B2025.12 | 26.3 | 37.8 | |
| CoTBackbone=LLaMA-3.1-8B-Instruct2026.02 | 23.5 | — | |
| Self-RAGBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs2026.03 | 22 | 34.38 | |
| FLAREBase LLM=Llama-3.1-8B-Instruct, Retriever=E5-base-v2, Top-k=3 paragraphs2026.03 | 21.6 | 24.35 | |
| IRCoTBackbone=Qwen-2.5-3B-Instruct2026.02 | 20 | — | |
| Direct InferenceBackbone=LLaMA-3.1-8B-Instruct2026.02 | 19.8 | — | |
| RAPTOR + Q3RType=Tree-RAG, Reranker=Qwen3-Reranker-8B (Q3R)2026.05 | 19.8 | 25.76 |