Question Answering on NQ (Natural Questions) (EM)
78.3EMInstructRetro 43B
Evaluation Results
| Method | Links | |
|---|---|---|
| InstructRetro 43BRetrieval-Augmented Generation=Enabled, Model Scale=43B2024.07 | 78.3 | |
| Raven 11BRetrieval-Augmented Generation=Enabled, Model Scale=11B2024.07 | 65.7 | |
| Llama3-RankRAG 70BRetrieval-Augmented Generation=Enabled, Zero-shot=true, Model Scale=70B, Backbone=Llama32024.07 | 54.2 | |
| DeepSeek-V2 236BRetrieval-Augmented Generation=Disabled, Few-shot=5-shot, Model Scale=236B2024.07 | 53.4 | |
| Llama3-RankRAG 8BRetrieval-Augmented Generation=Enabled, Zero-shot=true, Model Scale=8B, Backbone=Llama32024.07 | 50.6 | |
| BAR-RAGBackbone=LLaMA-3.1-8B-Instruct, Number of Iterations=32026.02 | 49.5 | |
| BAR-RAGBackbone=LLaMA-3.1-8B-Instruct, Number of Iterations=22026.02 | 49 | |
| BAR-RAGBackbone=LLaMA-3.1-8B-Instruct, Number of Iterations=12026.02 | 47.5 | |
| Llama3-ChatQA-1.5 70BRetrieval-Augmented Generation=Enabled, Model Scale=70B, Backbone=Llama32024.07 | 47 | |
| BAR-RAGBackbone=Qwen-2.5-7B-Instruct, Number of Iterations=32026.02 | 46.9 | |
| GPT-3.5-turbo-1106 RAGRetrieval-Augmented Generation=Enabled2024.07 | 46.7 | |
| SubSearch-base (Ours)Alg.=GRPO, Supervision=RL (SFT-Free)2026.04 | 46.3 | |
| BAR-RAGBackbone=Qwen-2.5-7B-Instruct, Number of Iterations=22026.02 | 46.1 | |
| RAG SFTBackbone=LLaMA-3.1-8B-Instruct2026.02 | 45.8 | |
| BAR-RAGBackbone=Qwen-2.5-7B-Instruct, Number of Iterations=12026.02 | 44.7 | |
| O2-Searcher-instructAlg.=GRPO, Supervision=SFT + RL (Upper Bound)2026.04 | 44.4 | |
| RAG w/RerankerBackbone=LLaMA-3.1-8B-Instruct2026.02 | 43.6 | |
| Llama3-Instruct 70BRetrieval-Augmented Generation=Enabled, Model Scale=70B, Backbone=Llama32024.07 | 42.7 | |
| RAG SFTBackbone=Qwen-2.5-7B-Instruct2026.02 | 42.7 | |
| RAGBackbone=LLaMA-3.1-8B-Instruct2026.02 | 42.7 | |
| Llama3-ChatQA-1.5 8BRetrieval-Augmented Generation=Enabled, Model Scale=8B, Backbone=Llama32024.07 | 42.4 | |
| Search-R1-baseAlg.=GRPO, Supervision=RL (SFT-Free)2026.04 | 42.1 | |
| InForage-instructAlg.=PPO, Supervision=SFT + RL (Upper Bound)2026.04 | 42.1 | |
| BAR-RAGBackbone=Qwen-2.5-3B-Instruct, Number of Iterations=32026.02 | 42 | |
| GPT-4-turbo-2024-0409Retrieval-Augmented Generation=Disabled, Zero-shot=true2024.07 | 41.5 | |
| BAR-RAGBackbone=Qwen-2.5-3B-Instruct, Number of Iterations=22026.02 | 41.5 | |
| ZeroSearch-instructAlg.=REINF., Supervision=RL (SFT-Free)2026.04 | 41.4 | |
| RAG w/RerankerBackbone=Qwen-2.5-7B-Instruct2026.02 | 40.5 | |
| GPT-4-0613 RAGRetrieval-Augmented Generation=Enabled2024.07 | 40.4 | |
| GPT-4-0613Retrieval-Augmented Generation=Disabled, Zero-shot=true2024.07 | 40.3 | |
| GPT-4-turbo-2024-0409 RAGRetrieval-Augmented Generation=Enabled2024.07 | 40.3 | |
| Mixtral-8x22B-InstructRetrieval-Augmented Generation=Disabled, Few-shot=5-shot, Model Scale=8x22B2024.07 | 40.1 | |
| BAR-RAGBackbone=Qwen-2.5-3B-Instruct, Number of Iterations=12026.02 | 40.1 | |
| Query Decomp-base (Ours)Alg.=GRPO, Supervision=RL (SFT-Free)2026.04 | 40.1 | |
| Search-R1-instructAlg.=GRPO, Supervision=RL (SFT-Free)2026.04 | 39.7 | |
| ReFeedRetrieval-Augmented Generation=Enabled, Backbone=InstructGPT/CodeX2024.07 | 39.6 | |
| ZeroSearch-baseAlg.=GRPO, Supervision=RL (SFT-Free)2026.04 | 39.4 | |
| RAGBackbone=Qwen-2.5-7B-Instruct2026.02 | 39.3 | |
| RAG SFTBackbone=Qwen-2.5-3B-Instruct2026.02 | 38.9 | |
| GPT-3.5-turbo-1106Retrieval-Augmented Generation=Disabled, Zero-shot=true2024.07 | 38.6 | |
| GLaM 64BRetrieval-Augmented Generation=Disabled, Zero-shot=true, Model Scale=64B2024.07 | 37.5 | |
| PaLM2 540BRetrieval-Augmented Generation=Disabled, Few-shot=5-shot, Model Scale=540B2024.07 | 37.1 | |
| RECOMP 20BRetrieval-Augmented Generation=Enabled, Model Scale=20B2024.07 | 37 | |
| RAG w/RerankerBackbone=Qwen-2.5-3B-Instruct2026.02 | 35.6 | |
| RA-DIT 65BRetrieval-Augmented Generation=Enabled, Model Scale=65B2024.07 | 35.2 | |
| RAGBackbone=Qwen-2.5-3B-Instruct2026.02 | 34.8 | |
| RAGAlg.=–, Supervision=Baselines2026.04 | 34.8 | |
| GenReadRetrieval-Augmented Generation=Enabled, Backbone=InstructGPT/CodeX2024.07 | 32.5 | |
| R-Search-instruct (r.)Alg.=GRPO, Supervision=RL (SFT-Free)2026.04 | 31.9 | |
| Retrieve-ReadRetrieval-Augmented Generation=Enabled, Backbone=InstructGPT/CodeX2024.07 | 31.7 | |
| Llama3-Instruct 8BRetrieval-Augmented Generation=Enabled, Model Scale=8B, Backbone=Llama32024.07 | 30.9 | |
| InstructGPTRetrieval-Augmented Generation=Disabled, Zero-shot=true2024.07 | 29.9 | |
| RePlug 65BRetrieval-Augmented Generation=Enabled, Model Scale=65B2024.07 | 28.8 | |
| CoTBackbone=LLaMA-3.1-8B-Instruct2026.02 | 27.8 | |
| Atlas 11BRetrieval-Augmented Generation=Enabled, Model Scale=11B2024.07 | 26.7 | |
| SFTAlg.=–, Supervision=Baselines2026.04 | 24.9 | |
| Search-o1Alg.=–, Supervision=Baselines2026.04 | 23.8 | |
| IRCoTBackbone=LLaMA-3.1-8B-Instruct2026.02 | 23.5 | |
| R1-baseAlg.=PPO, Supervision=RL (SFT-Free)2026.04 | 22.6 | |
| IRCoTBackbone=Qwen-2.5-7B-Instruct2026.02 | 22.4 | |
| PaLM2 540BRetrieval-Augmented Generation=Disabled, Zero-shot=true, Model Scale=540B2024.07 | 21.2 | |
| FLAN-LaMDA 137BRetrieval-Augmented Generation=Disabled, Zero-shot=true, Model Scale=137B2024.07 | 20.7 | |
| Direct InferenceBackbone=LLaMA-3.1-8B-Instruct2026.02 | 18.4 | |
| Direct InferenceBackbone=Qwen-2.5-7B-Instruct2026.02 | 13.4 | |
| IRCoTBackbone=Qwen-2.5-3B-Instruct2026.02 | 11.1 | |
| Direct InferenceBackbone=Qwen-2.5-3B-Instruct2026.02 | 10.6 | |
| Direct InferenceAlg.=–, Supervision=Baselines2026.04 | 10.6 | |
| CoTBackbone=Qwen-2.5-7B-Instruct2026.02 | 4.8 | |
| CoTBackbone=Qwen-2.5-3B-Instruct2026.02 | 2.3 | |
| CoTAlg.=–, Supervision=Baselines2026.04 | 2.3 |