Question Answering on Natural Questions (NQ) (test)
76Exact MatchSearch-R1 (EM)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Search-R1 (EM)Backbone=Qwen3-4B-Instruct-2507, Supervision Strategy=W. Gold Supervision2026.04 | 76 | — | |
| CCSBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=W.O. Gold Supervision2026.04 | 74.4 | — | |
| TTRLBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=W.O. Gold Supervision2026.04 | 74 | — | |
| CJBackbone=Qwen3-32B, Supervision Strategy=W.O. Gold Supervision2026.04 | 73.3 | — | |
| CCSBackbone=Qwen3-32B, Supervision Strategy=W.O. Gold Supervision2026.04 | 73.3 | — | |
| Search-o1Backbone=Qwen3-32B, Supervision Strategy=Model Inference2026.04 | 72.2 | — | |
| RAGBackbone=Qwen3-32B, Supervision Strategy=Model Inference2026.04 | 71.4 | — | |
| RLIFBackbone=Qwen3-32B, Supervision Strategy=W.O. Gold Supervision2026.04 | 71.4 | — | |
| CCSBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=W.O. Gold Supervision2026.04 | 71.2 | — | |
| CJBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=W.O. Gold Supervision2026.04 | 70.9 | — | |
| CJBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=W.O. Gold Supervision2026.04 | 70.4 | — | |
| TTRLBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=W.O. Gold Supervision2026.04 | 69.5 | — | |
| Search-R1 (EM)Backbone=Qwen2.5-7B-Instruct, Supervision Strategy=W. Gold Supervision2026.04 | 69.3 | — | |
| RAGBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=Model Inference2026.04 | 68.8 | — | |
| RLIFBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=W.O. Gold Supervision2026.04 | 68.6 | — | |
| Search-R1 (EM)Backbone=Qwen3-32B, Supervision Strategy=W. Gold Supervision2026.04 | 68.1 | — | |
| RLIFBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=W.O. Gold Supervision2026.04 | 67.2 | — | |
| Search-o1Backbone=Qwen2.5-7B-Instruct, Supervision Strategy=Model Inference2026.04 | 65.4 | — | |
| RAGBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=Model Inference2026.04 | 64.5 | — | |
| TTRLBackbone=Qwen3-32B, Supervision Strategy=W.O. Gold Supervision2026.04 | 64 | — | |
| IRCoTBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=Model Inference2026.04 | 61.3 | — | |
| IRCoTBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=Model Inference2026.04 | 60.8 | — | |
| ATLAS2022.12 | 60.4 | — | |
| Search-o1Backbone=Qwen3-4B-Instruct-2507, Supervision Strategy=Model Inference2026.04 | 54.1 | — | |
| FiDOConfig=L-XXL2022.12 | 53.2 | — | |
| IRCoTBackbone=Qwen3-32B, Supervision Strategy=Model Inference2026.04 | 52.2 | — | |
| EAR RD + FiDInput Passages to FiD=1002023.05 | 52.1 | — | |
| Liu et al. (2022) + FiDInput Passages to FiD=1002023.05 | 51.7 | — | |
| FiD-LSource=ours2022.12 | 51.5 | — | |
| FiD-LSource=Izacard and Grave, 20212022.12 | 51.4 | — | |
| DPR + FiDInput Passages to FiD=1002023.05 | 51.4 | — | |
| EAR RI + FiDInput Passages to FiD=1002023.05 | 51.4 | — | |
| SEAL + FiDInput Passages to FiD=1002023.05 | 50.7 | — | |
| GAR + FiDInput Passages to FiD=1002023.05 | 50.6 | — | |
| CoTBackbone=Qwen3-32B, Supervision Strategy=Model Inference2026.04 | 48.8 | — | |
| Direct InferenceBackbone=Qwen3-32B, Supervision Strategy=Model Inference2026.04 | 48.5 | — | |
| Grounded DecodingStrategy=adaptive2026.05 | 47.6 | 55.2 | |
| Grounded DecodingStrategy=static2026.05 | 45.8 | 53.5 | |
| RETRO2022.12 | 45.5 | — | |
| SFTBackbone=Qwen3-32B, Supervision Strategy=W. Gold Supervision2026.04 | 45.4 | — | |
| CoCoA2026.05 | 45.1 | 52.8 | |
| COIECD2026.05 | 44.8 | 52.5 | |
| RAG2022.12 | 44.5 | — | |
| RAGInput Passages to FiD=1002023.05 | 44.5 | — | |
| AdaCAD2026.05 | 44.3 | 52.1 | |
| CoTBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=Model Inference2026.04 | 43 | — | |
| DoLa2026.05 | 42.5 | 50.1 | |
| CAD2026.05 | 42.1 | 50.4 | |
| kNN-LM2026.05 | 41.8 | 49.8 | |
| DPR + ExtractiveInput Passages to FiD=1002023.05 | 41.5 | — | |
| Standard RAG2026.05 | 41.2 | 49.3 | |
| REALM2022.12 | 40.4 | — | |
| CoTBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=Model Inference2026.04 | 40.4 | — | |
| EAR RD + FiDInput Passages to FiD=102023.05 | 39.6 | — | |
| CriSPOModel=Claude Sonnet, Shots=64-shot2024.10 | 38.7 | — | |
| CriSPOModel=Claude Sonnet, Shots=0-shot2024.10 | 38.3 | — | |
| CriSPOModel=Claude Instant, Shots=64-shot2024.10 | 37.8 | — | |
| SFTBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=W. Gold Supervision2026.04 | 36.6 | — | |
| CriSPOModel=Claude Instant, Shots=0-shot2024.10 | 36.5 | — | |
| EAR RI + FiDInput Passages to FiD=102023.05 | 35.5 | — | |
| T5-XXL2022.12 | 35.2 | — | |
| Direct InferenceBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=Model Inference2026.04 | 34.9 | — | |
| Direct InferenceBackbone=Qwen2.5-7B-Instruct, Supervision Strategy=Model Inference2026.04 | 34.4 | — | |
| ManualModel=Claude Instant, Shots=0-shot2024.10 | 34 | — | |
| ManualModel=Claude Instant, Shots=64-shot2024.10 | 33.4 | — | |
| Self-RAGBackbone=LLaMA2-7B, Retriever=BM252025.04 | 32.3 | 40.2 | |
| ManualModel=Claude Sonnet, Shots=64-shot2024.10 | 32 | — | |
| SFTBackbone=Qwen3-4B-Instruct-2507, Supervision Strategy=W. Gold Supervision2026.04 | 31.2 | — | |
| GAR + FiDInput Passages to FiD=102023.05 | 30.5 | — | |
| ManualModel=Claude Sonnet, Shots=0-shot2024.10 | 26.6 | — | |
| DioRBackbone=LLaMA2-7B, Retriever=BM252025.04 | 26.2 | 35.9 | |
| FLAREBackbone=LLaMA2-7B, Retriever=BM252025.04 | 25.3 | 35.9 | |
| RaDIOBackbone=LLaMA2-7B, Retriever=BM252025.04 | 24.6 | 34 | |
| DRAGINBackbone=LLaMA2-7B, Retriever=BM252025.04 | 23.2 | 33.2 | |
| CoTBackbone=LLaMA2-7B, Retriever=BM252025.04 | 13.4 | 18.7 | |
| OPROModel=Claude Instant2024.10 | 8 | — | |
| OPROModel=Claude Sonnet2024.10 | 6.7 | — | |
| CoTMethod Category=Prompt-based2025.10 | — | 19.8 | |
| CoT+RAGMethod Category=Prompt-based2025.10 | — | 42 | |
| DeepResearcherMethod Category=Outcome-reward RL-based2025.10 | — | 39.6 | |
| GiGPOMethod Category=Step-reward RL-based2025.10 | — | 46.4 | |
| IGPOMethod Category=Step-reward RL-based2025.10 | — | 46.4 | |
| R1-searcherMethod Category=Outcome-reward RL-based2025.10 | — | 35.4 | |
| Search-o1Method Category=Prompt-based2025.10 | — | 32.4 | |
| Search-r1-baseMethod Category=Outcome-reward RL-based2025.10 | — | 45.4 | |
| Search-r1-instructMethod Category=Outcome-reward RL-based2025.10 | — | 33.1 |