Open-domain Question Answering on TriviaQA
76.1EMRA-ISF
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RA-ISFModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 76.1 | — | — | — | |
| RA-ISFModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 71.4 | — | — | — | |
| S2G-RAGReasoner=Llama3-8B, Retriever=E5, Retrieval Setting=Dense retrieval2026.04 | 71.1 | 78 | — | — | |
| Self-RAG13BModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 69.3 | — | — | — | |
| Least-to-mostModel=GPT-3.5, Retrieval Configuration=Without Retrieval2024.03 | 68.8 | — | — | — | |
| FiD-largeModel Type=Retrieval-augmented2022.10 | 67.6 | — | — | — | |
| SKRknnModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 67.5 | — | — | — | |
| CORE (3B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 10 docs (with compressor trained on top 5 docs), # tok=372025.08 | 67.36 | 74.04 | — | — | |
| DirectModel=GPT-3.5, Retrieval Configuration=Without Retrieval2024.03 | 67.3 | — | — | — | |
| IRCoTModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 66.8 | — | — | — | |
| CORE (3B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 5 docs, # tok=382025.08 | 66.5 | 73.06 | — | — | |
| Deepseek-V3 (671B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 10 docs (with compressor trained on top 5 docs), # tok=532025.08 | 65.29 | 74.45 | — | — | |
| Deepseek-V3 (671B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 5 docs, # tok=512025.08 | 65.28 | 74.33 | — | — | |
| FiD-baseModel Type=Retrieval-augmented, Backbone=T5-base2022.10 | 65 | — | — | — | |
| RAG-CriticReasoner=Llama3-8B, Retriever=E5, Retrieval Setting=Dense retrieval2026.04 | 65 | 75.9 | — | — | |
| Top10 DocumentsDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=None, Retrieval setting=Top10, # tok=14282025.08 | 64.4 | 72.92 | — | — | |
| RAGModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 64.2 | — | — | — | |
| Top5 DocumentsDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=None, Retrieval setting=Top5, # tok=7152025.08 | 64.1 | 72.48 | — | — | |
| Top3 DocumentsDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=None, Retrieval setting=Top3, # tok=4302025.08 | 62.6 | 71.02 | — | — | |
| RECOMP (3B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 10 docs (with compressor trained on top 5 docs), # tok=442025.08 | 62.05 | 69.73 | — | — | |
| RECOMP (3B)Downstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 5 docs, # tok=472025.08 | 61.83 | 69.2 | — | — | |
| M3POBackbone=Qwen2.5-3B-Instruct, Strategy=Multi-Path Perception Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 61 | — | — | — | |
| Top1 DocumentDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=None, Retrieval setting=Top1, # tok=1432025.08 | 60.82 | 68.7 | — | — | |
| HRPOBackbone=Qwen2.5-3B-Instruct, Strategy=Hybrid Reasoning Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 59.3 | — | — | — | |
| Standard RAGReasoner=Llama3-8B, Retriever=E5, Retrieval Setting=Dense retrieval2026.04 | 58.8 | 68.3 | — | — | |
| REPLUGModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 58.6 | — | — | — | |
| RAGBackbone=Qwen2.5-7B-Instruct, Strategy=Retrieval-Augmented Generation, Context Retrieval=Top-3 documents2025.12 | 58.5 | — | — | — | |
| DPRModel Type=Retrieval-augmented2022.10 | 57.9 | — | — | — | |
| llama3.2-3BDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 10 docs (with compressor trained on top 5 docs), # tok=572025.08 | 57.2 | 65.88 | — | — | |
| GRPOBackbone=Qwen2.5-3B-Instruct, Strategy=GRPO, Context Retrieval=Top-3 documents2025.12 | 57 | — | — | — | |
| RAGModel Type=Retrieval-augmented2022.10 | 56.8 | — | — | — | |
| M3POBackbone=Qwen2.5-1.5B-Instruct, Strategy=Multi-Path Perception Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 56.8 | — | — | — | |
| llama3.2-3BDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=Compression of top 5 docs, # tok=592025.08 | 56.5 | 65.21 | — | — | |
| PPOBackbone=Qwen2.5-3B-Instruct, Strategy=PPO, Context Retrieval=Top-3 documents2025.12 | 56.3 | — | — | — | |
| REALMModel Type=Retrieval-augmented2022.10 | 55.8 | — | — | — | |
| FLAREReasoner=Llama3-8B, Retriever=E5, Retrieval Setting=Dense retrieval2026.04 | 55.8 | 63.2 | — | — | |
| SKRknnModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 55.4 | — | — | — | |
| HRPOBackbone=Qwen2.5-1.5B-Instruct, Strategy=Hybrid Reasoning Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 55.3 | — | — | — | |
| RAGBackbone=Qwen2.5-3B-Instruct, Strategy=Retrieval-Augmented Generation, Context Retrieval=Top-3 documents2025.12 | 54.4 | — | — | — | |
| No RetrievalDownstream LLM=Qwen2.5-14B-Instruct, Compression setting=None, # tok=02025.08 | 53.23 | 59.98 | — | — | |
| PPOBackbone=Qwen2.5-1.5B-Instruct, Strategy=PPO, Context Retrieval=Top-3 documents2025.12 | 52.7 | — | — | — | |
| DensePhrasesModel Type=Retrieval-only2022.10 | 50.7 | — | — | — | |
| RePAQ rerankModel Type=Retrieval-augmented2022.10 | 48.9 | — | — | — | |
| IRCoTModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 48.3 | — | — | — | |
| QAMATModel Type=Retrieval-augmented2022.10 | 48 | — | — | — | |
| GRPOBackbone=Qwen2.5-1.5B-Instruct, Strategy=GRPO, Context Retrieval=Top-3 documents2025.12 | 48 | — | — | — | |
| IRCoTBackbone=Qwen2.5-7B-Instruct, Strategy=Interleaved retrieval with CoT, Context Retrieval=Top-3 documents2025.12 | 47.8 | — | — | — | |
| RAGBackbone=Qwen2.5-1.5B-Instruct, Strategy=Retrieval-Augmented Generation, Context Retrieval=Top-3 documents2025.12 | 47.7 | — | — | — | |
| RAGModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 47 | — | — | — | |
| Least-to-mostModel=Llama-2-13b, Retrieval Configuration=Without Retrieval2024.03 | 45.2 | — | — | — | |
| EMAT-FKSVAccess Delay Setting=Fast Key, Slow Value (FKSV), Backbone=T5-base2022.10 | 44.4 | — | — | — | |
| LLM-CASBackbone=LLaMA2-7B-CHAT2025.12 | 44.31 | — | — | — | |
| Search-o1Backbone=Qwen2.5-7B-Instruct, Strategy=Search-o1, Context Retrieval=Top-3 documents2025.12 | 44.3 | — | — | — | |
| EMAT-SKSVAccess Delay Setting=Slow Key, Slow Value (SKSV), Backbone=T5-base2022.10 | 43.7 | — | — | — | |
| SADIBackbone=LLaMA2-7B-CHAT2025.12 | 43.5 | — | — | — | |
| CAABackbone=LLaMA2-7B-CHAT2025.12 | 43.2 | — | — | — | |
| ITIBackbone=LLaMA2-7B-CHAT2025.12 | 42.8 | — | — | — | |
| T5-11BModel Type=Parametric2022.10 | 42.3 | — | — | — | |
| BaselineBackbone=LLaMA2-7B-CHAT2025.12 | 41.6 | — | — | — | |
| RePAQ-xlargeModel Type=Retrieval-only2022.10 | 41.3 | — | — | — | |
| QABackbone=Qwen2.5-7B-Instruct, Strategy=Direct inference, Context Retrieval=Top-3 documents2025.12 | 40.8 | — | — | — | |
| RePAQ-baseModel Type=Retrieval-only2022.10 | 39.7 | — | — | — | |
| FIXED_E=128Experts=1282026.05 | 39.11 | — | — | — | |
| Vanilla LMModel=Llama-2-13b, Retrieval Configuration=Without Retrieval2024.03 | 38.5 | — | — | — | |
| Self-RAGReasoner=Llama3-8B, Retriever=E5, Retrieval Setting=Dense retrieval2026.04 | 38.2 | 53.4 | — | — | |
| T5-3BModel Type=Parametric2022.10 | 35.1 | — | — | — | |
| EMO (Stage 5)E=64→1282026.05 | 33.77 | — | — | — | |
| T5-largeModel Type=Parametric2022.10 | 29.5 | — | — | — | |
| SFTBackbone=Qwen2.5-3B-Instruct, Strategy=Supervised Fine-Tuning, Context Retrieval=Top-3 documents2025.12 | 29.2 | — | — | — | |
| FIXED_E=32Experts=322026.05 | 27.52 | — | — | — | |
| BART-largeModel Type=Parametric2022.10 | 26.7 | — | — | — | |
| T5-baseModel Type=Parametric2022.10 | 24.4 | — | — | — | |
| FIXED_E=16Experts=162026.05 | 22.58 | — | — | — | |
| EMO (Stage 4)E=32→642026.05 | 22.21 | — | — | — | |
| SFTBackbone=Qwen2.5-1.5B-Instruct, Strategy=Supervised Fine-Tuning, Context Retrieval=Top-3 documents2025.12 | 19.3 | — | — | — | |
| CoTBackbone=Qwen2.5-7B-Instruct, Strategy=Chain-of-Thought, Context Retrieval=Top-3 documents2025.12 | 18.5 | — | — | — | |
| EMO (Stage 3)E=16→322026.05 | 17.76 | — | — | — | |
| EMO (Stage 1)E=82026.05 | 14.62 | — | — | — | |
| EMO (Stage 2)E=8→162026.05 | 13.98 | — | — | — | |
| DS(±)Training Data=Distant Supervision (Positive + Negative)2019.04 | 0.544 | 0.602 | 83.7 | — | |
| SRC -> DS(±)Training Strategy=Stage-wise, Order=SRC then DS(±)2019.04 | 0.537 | 0.593 | 83.7 | — | |
| SRC + DS(±)Training Strategy=Lumping (Source + DS±)2019.04 | 0.531 | 0.586 | 83.7 | — | |
| SRCTraining Data=Source (TriviaQA)2019.04 | 0.51 | 0.563 | 83.7 | — | |
| Evidence Agg.2019.04 | 0.506 | 0.573 | — | — | |
| DS(±) -> SRCTraining Strategy=Stage-wise, Order=DS(±) then SRC2019.04 | 0.498 | 0.559 | 83.7 | — | |
| DS-QA2019.04 | 0.487 | 0.563 | — | — | |
| DS(+)Training Data=Distant Supervision (Positive only)2019.04 | 0.482 | 0.536 | 83.7 | — | |
| R³2019.04 | 0.473 | 0.537 | — | — | |
| GPT-Jparameters=6B, Protocol=Zero-shot, External Knowledge=No2023.07 | — | 1.7 | — | — | |
| GPT-Jparameters=6B, Protocol=Zero-shot, External Knowledge=Yes2023.07 | — | 2.1 | — | — | |
| OPT-30bparameters=30B, Protocol=Zero-shot, External Knowledge=No2023.07 | — | 16.3 | — | — | |
| OPT-30bparameters=30B, Protocol=Zero-shot, External Knowledge=Yes2023.07 | — | 6.9 | — | — | |
| T5-3bparameters=3B, Protocol=Zero-shot, External Knowledge=No2023.07 | — | 8.3 | — | — | |
| T5-3bparameters=3B, Protocol=Zero-shot, External Knowledge=Yes2023.07 | — | 5.6 | — | — | |
| T5-baseparameters=220M, Protocol=Zero-shot, External Knowledge=No2023.07 | — | 8.9 | — | — | |
| T5-baseparameters=220M, Protocol=Zero-shot, External Knowledge=Yes2023.07 | — | 13.1 | — | — | |
| T5-largeparameters=770M, Protocol=Zero-shot, External Knowledge=No2023.07 | — | 8.3 | — | — | |
| T5-largeparameters=770M, Protocol=Zero-shot, External Knowledge=Yes2023.07 | — | 9 | — | — | |
| UnifiedQA-3bparameters=3B, Protocol=Transfer-learning, External Knowledge=No2023.07 | — | 18.6 | — | — | |
| UnifiedQA-3bparameters=3B, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | — | 80 | — | — |