Question Answering on TriviaQA
86.68AccuracyRankCoT
Evaluation Results
| Method | Links | |
|---|---|---|
| RankCoTBackbone=Qwen2.5-14B-Instruct2025.02 | 86.68 | |
| PaLM-2-Lshots=1-shot2023.07 | 86.1 | |
| RankCoTBackbone=MiniCPM3-4B2025.02 | 85.2 | |
| RankCoTBackbone=Llama3-8B-Instruct2025.02 | 85.18 | |
| Kimi-K2 Base# Shots=5-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 85.1 | |
| LLAMA 2shots=1-shot, size=70B2023.07 | 85 | |
| ASTUTE RAGLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 84.1 | |
| DeepSeek-V3.2 Exp Base# Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.01 | 83.9 | |
| RerankBackbone=Llama3-8B-Instruct2025.02 | 83.51 | |
| DeepSeek-V3.1 Base# Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.01 | 83.5 | |
| InstructRAGLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 83 | |
| No RefinementBackbone=Llama3-8B-Instruct2025.02 | 82.85 | |
| ASTUTE RAGLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 82.7 | |
| SummaryBackbone=Llama3-8B-Instruct2025.02 | 82.09 | |
| No RAGLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 82 | |
| ASTUTE RAGLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 81.6 | |
| CoTBackbone=Llama3-8B-Instruct2025.02 | 81.45 | |
| PaLMshots=1-shot2023.07 | 81.4 | |
| No RefinementBackbone=MiniCPM3-4B2025.02 | 80.91 | |
| USCLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 80.9 | |
| InstructRAGLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 80.6 | |
| InstructRAGLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 80.6 | |
| MiMo-V2-Flash Base# Shots=5-shot, # Activated Params=15B, # Total Params=309B2026.01 | 80.3 | |
| USCLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 80.2 | |
| No RAGLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 80.2 | |
| Self-RouteLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 79.9 | |
| No RAGLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 79.5 | |
| No RefinementBackbone=Qwen2.5-14B-Instruct2025.02 | 79.49 | |
| Self-RouteLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 78.8 | |
| RobustRAGLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 78.1 | |
| RobustRAGLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 77.7 | |
| Self-RouteLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 77.7 | |
| Mixtral-8x7B-v0.1Model size=13 ~ 20B, shot=4-shot2024.03 | 77.6 | |
| GenReadLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 77.4 | |
| RAGLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 77.4 | |
| RAGLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 76.7 | |
| USCLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 76.7 | |
| BSDETECTORLLM=GPT-3.5 Turbo, computational_overhead=10x2023.08 | 76 | |
| RAGLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 76 | |
| LinUCB+KLAlgorithm Family=Contextual Linear2026.02 | 75.9 | |
| ChatGPTRetrieval=No, Proprietary=true2023.10 | 74.3 | |
| GenReadLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 74.2 | |
| MAIN-RAG-Llama38BRetrieval=Training-free, Backbone=Llama3-8B2024.12 | 74.1 | |
| ASTUTE RAGLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 73.9 | |
| Reference AnswerLLM=GPT-3.5 Turbo, temperature=02023.08 | 73.5 | |
| Self-RouteLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 73.5 | |
| Llama38BRetrieval=Training-free2024.12 | 73.1 | |
| GenReadLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 73.1 | |
| RobustRAGLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 71.7 | |
| MAIN-RAG-Mistral7BRetrieval=Training-free, Backbone=Mistral-7B2024.12 | 71 | |
| BSDETECTORLLM=Text-Davinci-003, computational_overhead=10x2023.08 | 70.5 | |
| Mistral 7BModality=Pretrained2023.10 | 69.9 | |
| Reference AnswerLLM=Text-Davinci-003, temperature=02023.08 | 69.8 | |
| LLaMA 2 13BModality=Pretrained2023.10 | 69.6 | |
| Mistral7BRetrieval=Training-free2024.12 | 69.4 | |
| SELF-RAGScale=13B, Retrieval=Yes2023.10 | 69.3 | |
| SAILScale=7B, Retrieval=Yes2023.10 | 69.2 | |
| GPT-3.5Parameter Scale=API, Evaluation Setup=4-shot2024.03 | 69.1 | |
| Mistral-7B-v0.1Model size=< 7B, shot=4-shot2024.03 | 68.9 | |
| Llama27BRetrieval=Training-free2024.12 | 68.9 | |
| GenReadLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 68.9 | |
| SMOOTHIE-GLOBALBackbone=Llama-22024.12 | 68.7 | |
| BEST-ON-VALBackbone=Llama-22024.12 | 68.7 | |
| Llama38BRetrieval=None2024.12 | 68.4 | |
| No-RewriteAlgorithm Family=Base2026.02 | 68.2 | |
| RAG (Open-domain)Model setting=Finetuned2023.02 | 68 | |
| No RAGLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 67.8 | |
| RobustRAGLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 67.5 | |
| LESATraining Steps=3k2025.02 | 67.15 | |
| LESATraining Steps=6k2025.02 | 67.05 | |
| Mistral 2501 Base2026.02 | 67 | |
| AlpacaScale=13B, Retrieval=Yes2023.10 | 66.9 | |
| Alpaca13BRetrieval=Training-free2024.12 | 66.9 | |
| RAGLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 66.8 | |
| SELF-RAGScale=7B, Retrieval=Yes2023.10 | 66.4 | |
| Self-RAG7BRetrieval=Training-based2024.12 | 66.4 | |
| USCLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 66.1 | |
| GPT-3number of parameters=175B2023.02 | 65.9 | |
| Ret-ChatGPTRetrieval=Yes, Proprietary=true2023.10 | 65.7 | |
| LLAMA2Size=13B, Tokens=2T, Context=4K2024.04 | 65.1 | |
| AlpacaScale=7B, Retrieval=Yes2023.10 | 64.1 | |
| Alpaca7BRetrieval=Training-free2024.12 | 64.1 | |
| ChatGLM3-6B-BaseModel size=< 7B, shot=4-shot2024.03 | 63.9 | |
| LLAMA 2 7BModality=Pretrained2023.10 | 63.8 | |
| InternLM2-20B-BaseModel size=13 ~ 20B, shot=4-shot2024.03 | 63.7 | |
| LLaMA ProTraining Steps=6k2025.02 | 63.68 | |
| AMALIA-9B-DPOModel Category=Fully open models, Training=DPO, Instruction-tuned=true2026.03 | 63.5 | |
| GemmaSize=8B, Tokens=6T, Context=8K2024.04 | 63.4 | |
| Gemma 3-12BModel Category=Open weight models, Instruction-tuned=true2026.03 | 63.2 | |
| Qwen 2.5-7BModel Category=Open weight models, Instruction-tuned=true2026.03 | 63.2 | |
| Gemma 2-9BModel Category=Open weight models, Instruction-tuned=true2026.03 | 62.7 | |
| Qwen-14BModel size=13 ~ 20B, shot=4-shot2024.03 | 62.5 | |
| MistralSize=7B, Context=16K2024.04 | 62.5 | |
| LLaMA ProTraining Steps=3k2025.02 | 62.5 | |
| InstructRAGLLM=Mistral-Nemo (2407), 12B, Evaluation Setting=zero-shot2024.10 | 61.8 | |
| AlpacaScale=13B, Retrieval=No2023.10 | 61.3 | |
| SMOOTHIE-LOCALBackbone=Llama-22024.12 | 61.3 | |
| Alpaca13BRetrieval=None2024.12 | 61.3 | |
| Mistral-7BModel Category=Open weight models, Instruction-tuned=true2026.03 | 61.1 | |
| Llama2-13BModel size=13 ~ 20B, shot=4-shot2024.03 | 60.7 |