Question Answering on PopQA
68.4AccuracyLogicGaze
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LogicGazeBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=RAG2026.01 | 68.4 | — | |
| IterDRAGBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=RAG2026.01 | 66.5 | — | |
| RankRAGBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=RAG2026.01 | 66.1 | — | |
| Llama3-RankRAG 70BRetrieval-Augmented Generation=Enabled, Zero-shot=true2024.07 | 65.4 | 59.9 | |
| AutoRAGBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=RAG2026.01 | 65.3 | — | |
| RQ-RAGBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=RAG2026.01 | 64.2 | — | |
| Llama3-RankRAG 8BRetrieval-Augmented Generation=Enabled, Zero-shot=true2024.07 | 64.1 | 57.6 | |
| MAIN-RAG-Llama38BRetrieval=Training-free, Backbone=Llama3-8B2024.12 | 64 | — | |
| Self-RAGBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=RAG2026.01 | 62.9 | — | |
| Llama38BRetrieval=Training-free2024.12 | 61.8 | — | |
| GPT-4-0613 RAGRetrieval-Augmented Generation=Enabled2024.07 | 61.4 | 44.3 | |
| Llama3-ChatQA-1.5 8BRetrieval-Augmented Generation=Enabled2024.07 | 59.8 | 52.6 | |
| LogicGazeBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 59.5 | — | |
| MAIN-RAG-Mistral7BRetrieval=Training-free, Backbone=Mistral-7B2024.12 | 58.9 | — | |
| GPT-4-turbo-2024-0409 RAGRetrieval-Augmented Generation=Enabled2024.07 | 58.4 | 39.5 | |
| Llama3-ChatQA-1.5 70BRetrieval-Augmented Generation=Enabled2024.07 | 58.3 | 50.9 | |
| IterDRAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 58.3 | — | |
| RankRAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 58.2 | — | |
| GPT-3.5-turbo-1106 RAGRetrieval-Augmented Generation=Enabled2024.07 | 57 | 49.9 | |
| SFT+RetBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=SFT2026.01 | 56.9 | — | |
| Zero-Shot-Chat+RetBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=Chat2026.01 | 56.5 | — | |
| Llama3-Instruct 70BRetrieval-Augmented Generation=Enabled2024.07 | 56.4 | 45.3 | |
| SFTBackbone=Qwen2.5-7B, Retrieval Status=Without Retrieval, Evaluation Protocol=SFT2026.01 | 56 | — | |
| AutoRAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 56 | — | |
| SELF-RAGScale=13B, Retrieval=Yes2023.10 | 55.8 | — | |
| Llama3-Instruct 8BRetrieval-Augmented Generation=Enabled2024.07 | 55.8 | 34.9 | |
| Mistral7BRetrieval=Training-free2024.12 | 55.5 | — | |
| SELF-RAGScale=7B, Retrieval=Yes2023.10 | 54.9 | — | |
| Self-RAG 7BRetrieval-Augmented Generation=Enabled2024.07 | 54.9 | — | |
| Self-RAG7BRetrieval=Training-based2024.12 | 54.9 | — | |
| RQ-RAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 54.9 | — | |
| Zero-Shot-ChatBackbone=Qwen2.5-7B, Retrieval Status=Without Retrieval, Evaluation Protocol=Chat2026.01 | 54.2 | — | |
| SAIL-7BBackbone=Qwen2.5-7B, Retrieval Status=With Retrieval, Evaluation Protocol=RAG2026.01 | 53.3 | — | |
| Self-RAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 53.3 | — | |
| InstructGPT w/ AARANCESetting=Giant Setting: Over 70B Size, Protocol=Zero-shot, Parameters=175B2023.05 | 52 | — | |
| Ret-Llama2-ChatScale=13B, Retrieval=Yes, Proprietary=true2023.10 | 51.8 | — | |
| Ret-Llama2-chat13BData=Proprietary, Retrieval=Yes2024.12 | 51.8 | — | |
| Llama27BRetrieval=Training-free2024.12 | 50.9 | — | |
| Ret-ChatGPTRetrieval=Yes, Proprietary=true2023.10 | 50.8 | — | |
| LatentMemMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 50.16 | — | |
| LatentMemMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 49.9 | — | |
| LatentMemMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 49.4 | — | |
| LatentMemMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 49.34 | — | |
| GenerativeMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 49.09 | — | |
| G-MemoryMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 48.96 | — | |
| VoyagerMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 48.88 | — | |
| G-MemoryMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 48.8 | — | |
| Llama2-FTScale=7B, Retrieval=Yes, Fine-tuned=true2023.10 | 48.7 | — | |
| Llama2-FT7BRetrieval=Training-based2024.12 | 48.7 | — | |
| VoyagerMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 48.68 | — | |
| GenerativeMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 48.55 | — | |
| MetaGPTMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 48.42 | — | |
| OAgentMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 48.3 | — | |
| MetaGPTMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 47.92 | — | |
| OAgentMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 47.9 | — | |
| OAgentMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 47.9 | — | |
| MetaGPTMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 47.68 | — | |
| SFT+RetBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=SFT2026.01 | 47.6 | — | |
| MetaGPTMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 47.41 | — | |
| GenerativeMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 47.32 | — | |
| G-MemoryMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 47.32 | — | |
| OAgentMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 47.25 | — | |
| G-MemoryMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 47.24 | — | |
| VoyagerMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 47.11 | — | |
| VoyagerMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 46.75 | — | |
| AlpacaScale=7B, Retrieval=Yes2023.10 | 46.7 | — | |
| Alpaca7BRetrieval=Training-free2024.12 | 46.7 | — | |
| SFTBackbone=LLaMA3-8B, Retrieval Setting=No, Evaluation Protocol=SFT2026.01 | 46.5 | — | |
| AlpacaScale=13B, Retrieval=Yes2023.10 | 46.1 | — | |
| Alpaca13BRetrieval=Training-free2024.12 | 46.1 | — | |
| GenerativeMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 46 | — | |
| Llama2Scale=13B, Retrieval=Yes2023.10 | 45.7 | — | |
| Llama213BRetrieval=Training-free2024.12 | 45.7 | — | |
| ReFeedRetrieval-Augmented Generation=Enabled2024.07 | 45.1 | 41.5 | |
| GPT-4oBackbone=GPT-4o, Retrieval Status=With Retrieval2026.01 | 45.1 | — | |
| RankCoTBackbone=Qwen2.5-14B-Instruct2025.02 | 44.45 | — | |
| ASTUTE RAGLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 44.4 | — | |
| InstructGPT w/ AARContrieverSetting=Giant Setting: Over 70B Size, Protocol=Zero-shot, Parameters=175B2023.05 | 43.9 | — | |
| Zero-Shot-Instruct+RetBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=Zero-Shot-Instruct2026.01 | 43.6 | — | |
| InstructGPT w/ ARSetting=Giant Setting: Over 70B Size, Protocol=Zero-shot, Parameters=175B2023.05 | 43.3 | — | |
| SAIL-7BBackbone=SAIL-7B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 42.7 | — | |
| ASTUTE RAGLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 42.1 | — | |
| Zero-Shot-InstructBackbone=LLaMA3-8B, Retrieval Setting=No, Evaluation Protocol=Zero-Shot-Instruct2026.01 | 41.3 | — | |
| No-memoryMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 41.2 | — | |
| RankCoTBackbone=Llama3-8B-Instruct2025.02 | 41.17 | — | |
| InstructRAGLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 41 | — | |
| Self-RouteLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 41 | — | |
| No-memoryMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 40.59 | — | |
| ASTUTE RAGLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 40.5 | — | |
| No-memoryMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 40.5 | — | |
| No-memoryMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 40.44 | — | |
| Flan-T5Large w/ AARANCESetting=Large Setting: T5 Large Size, Protocol=Zero-shot, Parameters=780M2023.05 | 39.3 | — | |
| Llama2Scale=7B, Retrieval=Yes2023.10 | 38.2 | — | |
| Self-RouteLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 38.2 | — | |
| Self-RouteLLM=Mistral-Large (2407), 128B, Evaluation Setting=zero-shot2024.10 | 38.2 | — | |
| Flan-T5XL w/ AARANCESetting=XL Setting: T5 XL Size, Protocol=Zero-shot, Parameters=3B2023.05 | 38 | — | |
| Flan-T5Base w/ AARANCESetting=Base Setting: T5 Base Size, Protocol=Zero-shot, Parameters=250M2023.05 | 37.7 | — | |
| USCLLM=Claude 3.5 Sonnet (20240620), Evaluation Setting=zero-shot2024.10 | 37.6 | — | |
| USCLLM=Gemini 1.5 Pro (002), Evaluation Setting=zero-shot2024.10 | 37.6 | — | |
| OLMo 2 32BParameters=32B2025.12 | 37.2 | — |