Question Answering on WebQuestions (WebQs)
63.5AccuracyGemini2.5-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini2.5-ProEvaluation Protocol=Closed-Source2026.01 | 63.5 | |
| GPT-4oEvaluation Protocol=Closed-Source2026.01 | 63 | |
| GPT-5Evaluation Protocol=Closed-Source2026.01 | 61.5 | |
| GPT-4.1Evaluation Protocol=Closed-Source2026.01 | 61.5 | |
| ATLAS (cluster)Evaluation Protocol=In-Distribution2026.01 | 53.6 | |
| CDKCBackbone=Qwen2.5-14B-Instruct2026.02 | 52.51 | |
| ATLAS (RL)Evaluation Protocol=Out-of-Distribution2026.01 | 52.2 | |
| BertRouterEvaluation Protocol=Out-of-Distribution2026.01 | 51.4 | |
| ATLAS (cluster)Evaluation Protocol=Out-of-Distribution2026.01 | 51.4 | |
| GRPOBackbone=Qwen2.5-14B-Instruct2026.02 | 51.23 | |
| RouterDCEvaluation Protocol=Out-of-Distribution2026.01 | 50.8 | |
| BertRouterEvaluation Protocol=In-Distribution2026.01 | 50.4 | |
| GradOTSetting=Plug-in, Backbone=LLaMA-13B2025.07 | 48.1 | |
| RouterDCEvaluation Protocol=In-Distribution2026.01 | 47.6 | |
| RAGSetting=Fine-tuned, QA Environment=Open-Domain2020.05 | 45.5 | |
| T5-11B+SSMSetting=Fine-tuned, QA Environment=Closed-Book, Auxiliary Training=SSM2020.05 | 44.7 | |
| MLPRouterEvaluation Protocol=Out-of-Distribution2026.01 | 43.7 | |
| CDKCBackbone=Qwen2.5-3B-Instruct2026.02 | 43.26 | |
| GRPOBackbone=Qwen2.5-3B-Instruct2026.02 | 42.96 | |
| CGKEBackbone=Qwen2.5-14B-Instruct2026.02 | 41.93 | |
| GPT-3Learning Protocol=Few-Shot, QA Environment=Closed-Book2020.05 | 41.5 | |
| MLPRouterEvaluation Protocol=In-Distribution2026.01 | 40.4 | |
| ZS RouterEvaluation Protocol=Training-free, Prompting Strategy=Zero-shot2026.01 | 39.2 | |
| CoTBackbone=Qwen2.5-14B-Instruct2026.02 | 38.83 | |
| EmulatorSetting=Fine-tuned, Backbone=LLaMA-13B2025.07 | 37.9 | |
| T5-11BSetting=Fine-tuned, QA Environment=Closed-Book2020.05 | 37.4 | |
| Vanilla SFTBackbone=Qwen2.5-14B-Instruct2026.02 | 35.97 | |
| FS RouterEvaluation Protocol=Training-free, Prompting Strategy=Few-shot2026.01 | 35.8 | |
| OT†Setting=Plug-in, Backbone=LLaMA-13B2025.07 | 35.4 | |
| CGKEBackbone=Qwen2.5-3B-Instruct2026.02 | 34.79 | |
| RAGBackbone=Qwen2.5-14B-Instruct2026.02 | 34.35 | |
| Vanilla LLMBackbone=Qwen2.5-14B-Instruct2026.02 | 33.46 | |
| Random RouterEvaluation Protocol=Training-free2026.01 | 32.1 | |
| Full LLMSetting=Fine-tuning (FT), Backbone=OPT-1.3B2025.07 | 31.2 | |
| CoTBackbone=Qwen2.5-3B-Instruct2026.02 | 31.2 | |
| RAGBackbone=Qwen2.5-3B-Instruct2026.02 | 30.51 | |
| GradOTSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 29.2 | |
| GPT-3number of parameters=175B2023.02 | 29 | |
| Vanilla SFTBackbone=Qwen2.5-3B-Instruct2026.02 | 28.74 | |
| GradOTSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 28.4 | |
| ScaleOTSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 28.2 | |
| Vanilla LLMBackbone=Qwen2.5-3B-Instruct2026.02 | 26.38 | |
| ToolformerAPI calls=enabled2023.02 | 26.3 | |
| OTSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 26.2 | |
| ScaleOTSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 25.3 | |
| GPT-3Learning Protocol=One-Shot, QA Environment=Closed-Book2020.05 | 25.3 | |
| OTSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 24.3 | |
| CRaShSetting=Plug-in, Backbone=OPT-1.3B2025.07 | 23.7 | |
| OT†Setting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 21.8 | |
| CRaShSetting=Emulator Fine-tuning (Emu. FT), Backbone=OPT-1.3B2025.07 | 21.8 | |
| OT†Setting=Plug-in, Backbone=OPT-1.3B2025.07 | 21.4 | |
| GPT-3 (Original)few-shot (k)=642021.08 | 19.6 | |
| GPT-3 (SLW 8x Bsz)few-shot (k)=642021.08 | 19.4 | |
| ToolformerAPI calls=disabled2023.02 | 18.9 | |
| OPTnumber of parameters=66B2023.02 | 18.6 | |
| GPT-J2023.02 | 18.5 | |
| GPT-3 (Baseline repro)few-shot (k)=642021.08 | 18.4 | |
| GPT-J + CCpre-training=CCNet2023.02 | 18.4 | |
| GPT-3Learning Protocol=Zero-Shot, QA Environment=Closed-Book2020.05 | 14.4 | |
| CRaShSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 4.7 | |
| Full LLMSetting=Zero-shot (ZS), Backbone=OPT-1.3B2025.07 | 4.6 | |
| OTSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 1.3 | |
| ScaleOTSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 0.1 | |
| OT†Setting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 0 | |
| GradOTSetting=Emulator Zero-shot (Emu. ZS), Backbone=OPT-1.3B2025.07 | 0 | |
| Full ModelSetting=Zero-shot, Backbone=LLaMA-13B2025.07 | 0 | |
| EmulatorSetting=Zero-shot, Backbone=LLaMA-13B2025.07 | 0 |