Open-Domain Question-Answering on NQ
61.6AccuracyRetroLLM (Ours)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| RetroLLM (Ours)Generation Type=Retrieval within Generation, Backbone=Mistral-7B-Instruct2024.12 | 61.6 | — | 49.8 | 302 | — | |
| HiPRAGBackbone=Qwen2.5-7B-Instruct2026.05 | 56 | — | — | — | 1.96 | |
| HiPRAGBackbone=Qwen2.5-3B-Instruct2026.05 | 52.5 | — | — | — | 1.39 | |
| Naive RAGGeneration Type=Retrieval-augmented Generation, Retriever=E5-base-en2024.12 | 52.4 | — | 41.1 | 919 | — | |
| Iter-RetGenGeneration Type=Retrieval-augmented Generation, Retriever=E5-base-en2024.12 | 51.7 | — | 48.4 | 3,002 | — | |
| Adaptive-RAGGeneration Type=Retrieval-augmented Generation, Retriever=E5-base-en2024.12 | 50.5 | — | 46.6 | 946 | — | |
| IRCoTGeneration Type=Retrieval-augmented Generation, Retriever=E5-base-en2024.12 | 49.6 | — | 45.9 | 1,598 | — | |
| HippoRAG + QREAM-FTRAG Framework=HippoRAG Integration2026.04 | 49.5 | — | 43.9 | — | — | |
| HippoRAG + QREAM-ICLRAG Framework=HippoRAG Integration2026.04 | 48.8 | — | 43.1 | — | — | |
| HippoRAG + FaviCompRAG Framework=HippoRAG Integration2026.04 | 48.2 | — | 42.8 | — | — | |
| SAASBackbone=Qwen2.5-7B-Instruct2026.05 | 47.8 | — | — | — | 0.74 | |
| StepSearchBackbone=Qwen2.5-7B-Instruct2026.05 | 47.7 | — | — | — | 1.17 | |
| HippoRAGRAG Framework=HippoRAG Integration2026.04 | 47 | — | 41.5 | — | — | |
| RFTBackbone=Qwen2.5-7B-Instruct2026.05 | 46.2 | — | — | — | 1.19 | |
| Search-R1Backbone=Qwen2.5-7B-Instruct2026.05 | 45.6 | — | — | — | 1.24 | |
| Self-RAG + QREAM-FTRAG Framework=Self-RAG Integration2026.04 | 45.3 | — | 44.9 | — | — | |
| Self-RAG + QREAM-ICLRAG Framework=Self-RAG Integration2026.04 | 45 | — | 44.2 | — | — | |
| Search-R1Backbone=Qwen2.5-3B-Instruct2026.05 | 45 | — | — | — | 1.2 | |
| QREAM-FTRAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 44.6 | — | 48 | — | — | |
| Self-RAG + FaviCompRAG Framework=Self-RAG Integration2026.04 | 44.1 | — | 42.5 | — | — | |
| QREAM-ICLRAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 43.8 | — | 47.5 | — | — | |
| SAASBackbone=Qwen2.5-3B-Instruct2026.05 | 43.6 | — | — | — | 0.72 | |
| Self-RAGRAG Framework=Self-RAG Integration2026.04 | 43.2 | — | 42.7 | — | — | |
| Retrieved Doc.RAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 42.6 | — | 47.1 | — | — | |
| CompActRAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 42.3 | — | 46.1 | — | — | |
| FaviCompRAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 42.3 | — | 46.6 | — | — | |
| Self-RAGGeneration Type=Retrieval-augmented Generation, Retriever=E5-base-en2024.12 | 41.8 | — | 45.2 | 1,203 | — | |
| Retrieved + GeneratedRAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 41.8 | — | 22.4 | — | — | |
| REPLUGGeneration Type=Retrieval-augmented Generation, Retriever=E5-base-en2024.12 | 41.6 | — | 41.2 | 903 | — | |
| RECOMPRAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 41.5 | — | 45.8 | — | — | |
| QREAM-FTRAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 41.5 | — | 40.1 | — | — | |
| StepSearchBackbone=Qwen2.5-3B-Instruct2026.05 | 41.2 | — | — | — | 1.37 | |
| QREAM-ICLRAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 41.1 | — | 40.5 | — | — | |
| RFTBackbone=Qwen2.5-3B-Instruct2026.05 | 40.6 | — | — | — | 1.2 | |
| FaviCompRAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 40.3 | — | 40.4 | — | — | |
| Retrieved Doc.RAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 40.2 | — | 39.3 | — | — | |
| Retrieved + GeneratedRAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 39.5 | — | 19.3 | — | — | |
| Generated Doc.RAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 39.1 | — | 23.1 | — | — | |
| RECOMPRAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 39.1 | — | 39.5 | — | — | |
| CompActRAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 38.8 | — | 38.9 | — | — | |
| Generated Doc.RAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 37.5 | — | 20.2 | — | — | |
| LongLLMLinguaRAG Framework=Standard RAG, Backbone LLM=Llama-3-8B-Instruct2026.04 | 35.4 | — | 40.9 | — | — | |
| LongLLMLinguaRAG Framework=Standard RAG, Backbone LLM=Mistral-7B-Instruct2026.04 | 34.3 | — | 36.4 | — | — | |
| Mistral-7BGeneration Type=Direct Generation2024.12 | 30.4 | — | 25.2 | 57 | — | |
| Direct InferenceBackbone=Qwen2.5-7B-Instruct2026.05 | 28.2 | — | — | — | — | |
| EMoEBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=EMoE, Inference Budget (k')=62025.09 | 27.87 | — | — | — | — | |
| EMoEBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=EMoE, Inference Budget (k')=42025.09 | 27.78 | — | — | — | — | |
| Llama3-8BGeneration Type=Direct Generation2024.12 | 27.6 | — | 30.1 | 50 | — | |
| Top-kBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=Top-k, Inference Budget (k')=62025.09 | 27.17 | — | — | — | — | |
| Top-kBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=Top-k, Inference Budget (k')=42025.09 | 26.91 | — | — | — | — | |
| Top-kBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=12 + 22025.09 | 26.81 | — | — | — | — | |
| EMoEBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=EMoE, Inference Budget (k')=22025.09 | 26.62 | — | — | — | — | |
| Top-kBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=6 + 22025.09 | 26.43 | — | — | — | — | |
| Top-kBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=3 + 22025.09 | 26.18 | — | — | — | — | |
| EMoEBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=12 + 22025.09 | 26.15 | — | — | — | — | |
| EMoEBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=6 + 22025.09 | 25.62 | — | — | — | — | |
| Top-kBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=Top-k, Inference Budget (k')=22025.09 | 25.55 | — | — | — | — | |
| EMoEBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=EMoE, Inference Budget (k')=12025.09 | 25.46 | — | — | — | — | |
| EMoEBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=3 + 22025.09 | 25.37 | — | — | — | — | |
| Top-kBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=Top-k, Inference Budget (k')=12025.09 | 22.89 | — | — | — | — | |
| EMoEBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=12 + 22025.09 | 22.8 | — | — | — | — | |
| Qwen2.5-7BGeneration Type=Direct Generation2024.12 | 21.8 | — | 21.3 | 52 | — | |
| EMoEBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=6 + 22025.09 | 21.72 | — | — | — | — | |
| Top-kBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=6 + 22025.09 | 21.63 | — | — | — | — | |
| Direct InferenceBackbone=Qwen2.5-3B-Instruct2026.05 | 21.6 | — | — | — | — | |
| Top-kBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=12 + 22025.09 | 21.05 | — | — | — | — | |
| EMoEBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=3 + 22025.09 | 20.44 | — | — | — | — | |
| EMoEBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=EMoE, Inference Budget (k')=162025.09 | 19.31 | — | — | — | — | |
| Top-kBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=3 + 22025.09 | 18.89 | — | — | — | — | |
| EMoEBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=EMoE, Inference Budget (k')=82025.09 | 18.67 | — | — | — | — | |
| Top-kBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=Top-k, Inference Budget (k')=162025.09 | 18.12 | — | — | — | — | |
| Top-kBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=Top-k, Inference Budget (k')=82025.09 | 17.95 | — | — | — | — | |
| Top-kBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=Top-k, Inference Budget (k')=42025.09 | 14.57 | — | — | — | — | |
| EMoEBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=EMoE, Inference Budget (k')=42025.09 | 14.52 | — | — | — | — | |
| AdaCompRetrieval Strategy=Adaptive Compression, Generator=LLAMA2-7B2024.09 | — | 40.13 | 70.96 | 441 | — | |
| BM25k=1, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 25.3 | 36.6 | 54 | — | |
| BM25k=2, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 26.7 | 38.6 | 90 | — | |
| BM25k=5, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 28.1 | 40.4 | 197 | — | |
| CORE (3B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 5 docs, # tok=322025.08 | — | 43.1 | 52.34 | — | — | |
| CORE (3B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 10 docs (with compressor trained on top 5 docs), # tok=332025.08 | — | 45.26 | 54.67 | — | — | |
| Deepseek-V3 (671B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 5 docs, # tok=542025.08 | — | 37.73 | 50.39 | — | — | |
| Deepseek-V3 (671B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 10 docs (with compressor trained on top 5 docs), # tok=562025.08 | — | 37.79 | 51.07 | — | — | |
| Densek=1, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 28.2 | 40.5 | 58 | — | |
| Densek=2, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 30.6 | 43.5 | 96 | — | |
| Densek=5, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 33.1 | 46.5 | 210 | — | |
| DPPBackbone=Llama-2-7B-Chat, Noise Type=Relevant Noise, Noise Ratio=60%2024.05 | — | 7.89 | — | — | — | |
| DPP + LPRBackbone=Llama-2-7B-Chat, Noise Type=Relevant Noise, Noise Ratio=60%2024.05 | — | 11.94 | — | — | — | |
| EJ-Full2025.12 | — | 42.3 | — | — | — | |
| FILCORetrieval Strategy=Retrieval with Compression, Generator=LLAMA2-7B2024.09 | — | 32.43 | 64.78 | 46 | — | |
| Full ContextGenerator=Meta-Llama-3-8B-Instruct, Protocol=Zero-shot cross-domain evaluation2026.06 | — | 0 | 1.1 | 23,588 | — | |
| JSA-RAG2025.08 | — | 51.05 | — | — | — | |
| Latent Memoryk=1, Memory tokens=1, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 22.2 | 33.7 | 26 | — | |
| Latent Memoryk=2, Memory tokens=1, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 22.7 | 34.6 | 35 | — | |
| Latent Memoryk=5, Memory tokens=1, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 23.4 | 35 | 62 | — | |
| Latent Memoryk=1, Memory tokens=8, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 25.5 | 37.2 | 33 | — | |
| Latent Memoryk=2, Memory tokens=8, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 24.9 | 36.8 | 49 | — | |
| Latent Memoryk=5, Memory tokens=8, Generator=Meta-Llama-3-8B-Instruct2026.06 | — | 26.7 | 38.5 | 97 | — | |
| LLAMA2-7BRetrieval Strategy=No Retrieval, Generator=LLAMA2-7B2024.09 | — | 26.98 | 62.51 | — | — | |
| llama3.2-3BDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 5 docs, # tok=602025.08 | — | 32.52 | 43.34 | — | — | |
| llama3.2-3BDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 10 docs (with compressor trained on top 5 docs), # tok=612025.08 | — | 33.18 | 43.59 | — | — |