Open Domain Question Answering on Natural Questions (NQ)
60.7Exact Match (EM)DeepSeek-v3
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-v3Category=Candidate Models2026.03 | 60.7 | |
| DeepSeek-R1Category=Candidate Models2026.03 | 60.3 | |
| Claude-Sonnet-4.5Category=Candidate Models2026.03 | 59.4 | |
| FineRouterCategory=Routers2026.03 | 59 | |
| Llama-3.3-70BCategory=Candidate Models2026.03 | 57.6 | |
| IPRCategory=Routers2026.03 | 57.6 | |
| kNNCategory=Routers2026.03 | 56.2 | |
| Mistral-LargeCategory=Candidate Models2026.03 | 56 | |
| Qwen3-235B-A22BCategory=Candidate Models2026.03 | 55.1 | |
| RouteLLMCategory=Routers2026.03 | 55 | |
| Llama-4-MaverickCategory=Candidate Models2026.03 | 53.8 | |
| MLPCategory=Routers2026.03 | 53.6 | |
| Fusion-in-DecoderModel scale=large2020.07 | 51.4 | |
| FIDReader Type=Generative2020.09 | 51.4 | |
| Mistral-SmallCategory=Candidate Models2026.03 | 50.3 | |
| RouterDCCategory=Routers2026.03 | 50.3 | |
| GPT-OSS-120BCategory=Candidate Models2026.03 | 48.3 | |
| GraphRouterCategory=Routers2026.03 | 48.3 | |
| Fusion-in-DecoderModel scale=base2020.07 | 48.2 | |
| Claude-Haiku-4.5Category=Candidate Models2026.03 | 46.6 | |
| GAR+Reader Type=Generative2020.09 | 45.3 | |
| RAG2020.07 | 44.5 | |
| RAGReader Type=Generative2020.09 | 44.5 | |
| M3POBackbone=Qwen2.5-3B-Instruct, Strategy=Multi-Path Perception Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 44.1 | |
| GAR+Reader Type=Extractive2020.09 | 43.8 | |
| SpanSeqGen2020.07 | 42.5 | |
| SpanSeqGenReader Type=Generative2020.09 | 42.2 | |
| GARReader Type=Extractive2020.09 | 41.8 | |
| DPR2020.07 | 41.5 | |
| DPRReader Type=Extractive2020.09 | 41.5 | |
| M3POBackbone=Qwen2.5-1.5B-Instruct, Strategy=Multi-Path Perception Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 41.4 | |
| REALM2020.07 | 40.4 | |
| REALMReader Type=Extractive2020.09 | 40.4 | |
| Qwen3-32BCategory=Candidate Models2026.03 | 39.5 | |
| GARReader Type=Generative2020.09 | 38.1 | |
| GRPOBackbone=Qwen2.5-3B-Instruct, Strategy=GRPO, Context Retrieval=Top-3 documents2025.12 | 38.1 | |
| HRPOBackbone=Qwen2.5-3B-Instruct, Strategy=Hybrid Reasoning Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 37.8 | |
| BM25 (ours)Reader Type=Extractive2020.09 | 37.7 | |
| T52020.07 | 36.6 | |
| T5Reader Type=Generative2020.09 | 36.6 | |
| HRPOBackbone=Qwen2.5-1.5B-Instruct, Strategy=Hybrid Reasoning Policy Optimization, Context Retrieval=Top-3 documents2025.12 | 36.4 | |
| PPOBackbone=Qwen2.5-3B-Instruct, Strategy=PPO, Context Retrieval=Top-3 documents2025.12 | 35.6 | |
| BM25 (ours)Reader Type=Generative2020.09 | 35.3 | |
| RAGBackbone=Qwen2.5-7B-Instruct, Strategy=Retrieval-Augmented Generation, Context Retrieval=Top-3 documents2025.12 | 34.9 | |
| RAGBackbone=Qwen2.5-3B-Instruct, Strategy=Retrieval-Augmented Generation, Context Retrieval=Top-3 documents2025.12 | 34.8 | |
| Graph Retriever2020.07 | 34.7 | |
| Graph RetrieverReader Type=Extractive2020.09 | 34.5 | |
| ORQAReader Type=Extractive2020.09 | 33.3 | |
| PPOBackbone=Qwen2.5-1.5B-Instruct, Strategy=PPO, Context Retrieval=Top-3 documents2025.12 | 32.7 | |
| Path RetrieverReader Type=Extractive2020.09 | 32.6 | |
| Path Retriever2020.07 | 31.7 | |
| ORQA2020.07 | 31.3 | |
| GPT-3shot=few-shot2020.07 | 29.9 | |
| GPT-3Reader Type=Generative2020.09 | 29.9 | |
| GRPOBackbone=Qwen2.5-1.5B-Instruct, Strategy=GRPO, Context Retrieval=Top-3 documents2025.12 | 29.3 | |
| Hard EM2020.07 | 28.8 | |
| RAGBackbone=Qwen2.5-1.5B-Instruct, Strategy=Retrieval-Augmented Generation, Context Retrieval=Top-3 documents2025.12 | 28.8 | |
| Hard EMReader Type=Extractive2020.09 | 28.1 | |
| SFTBackbone=Qwen2.5-3B-Instruct, Strategy=Supervised Fine-Tuning, Context Retrieval=Top-3 documents2025.12 | 24.9 | |
| IRCoTBackbone=Qwen2.5-7B-Instruct, Strategy=Interleaved retrieval with CoT, Context Retrieval=Top-3 documents2025.12 | 22.4 | |
| FIXED_E=128Experts=1282026.05 | 16.34 | |
| Search-o1Backbone=Qwen2.5-7B-Instruct, Strategy=Search-o1, Context Retrieval=Top-3 documents2025.12 | 15.1 | |
| QABackbone=Qwen2.5-7B-Instruct, Strategy=Direct inference, Context Retrieval=Top-3 documents2025.12 | 13.4 | |
| EMO (Stage 5)E=64→1282026.05 | 13.35 | |
| FIXED_E=32Experts=322026.05 | 11.86 | |
| MobileLLM-Flash 1.4BParameter Count=1.4B, Evaluation Protocol=5-shot2026.03 | 11.83 | |
| Gemma3 1BParameter Count=1B, Evaluation Protocol=5-shot2026.03 | 9.48 | |
| SFTBackbone=Qwen2.5-1.5B-Instruct, Strategy=Supervised Fine-Tuning, Context Retrieval=Top-3 documents2025.12 | 9.4 | |
| FIXED_E=16Experts=162026.05 | 9.11 | |
| EMO (Stage 4)E=32→642026.05 | 8.64 | |
| LFM2 1.2BParameter Count=1.2B, Evaluation Protocol=5-shot2026.03 | 8.6 | |
| MobileLLM-Flash 650MParameter Count=650M, Evaluation Protocol=5-shot2026.03 | 7.56 | |
| EMO (Stage 3)E=16→322026.05 | 7.23 | |
| LFM2 700MParameter Count=700M, Evaluation Protocol=5-shot2026.03 | 6.9 | |
| MobileLLM-Flash 350MParameter Count=350M, Evaluation Protocol=5-shot2026.03 | 5.9 | |
| EMO (Stage 1)E=82026.05 | 5.65 | |
| Llama3.2 1BParameter Count=1B, Evaluation Protocol=5-shot2026.03 | 5.48 | |
| EMO (Stage 2)E=8→162026.05 | 5.24 | |
| LFM2 350MParameter Count=350M, Evaluation Protocol=5-shot2026.03 | 4.96 | |
| CoTBackbone=Qwen2.5-7B-Instruct, Strategy=Chain-of-Thought, Context Retrieval=Top-3 documents2025.12 | 4.8 | |
| Qwen3 0.6BParameter Count=0.6B, Evaluation Protocol=5-shot2026.03 | 4.32 | |
| Gemma3 270MParameter Count=270M, Evaluation Protocol=5-shot2026.03 | 4.04 |