Question Answering on TriviaQA (EM and Accuracy)
94.5AccuracyGPT-4-0613
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4-0613Retrieval-Augmented Generation=Disabled2024.07 | 94.5 | 84.8 | |
| GPT-4-turbo-2024-0409Retrieval-Augmented Generation=Disabled2024.07 | 94.3 | 80 | |
| OSCAR-llamaBackbone=Mistral-24B2025.03 | 93 | — | |
| Llama3-RankRAG 70BRetrieval-Augmented Generation=Enabled, Zero-shot=true2024.07 | 92.3 | 86.5 | |
| No compressionBackbone=Mistral-7B2025.03 | 92 | — | |
| RECOMPBackbone=Mistral-7B2025.03 | 92 | — | |
| ProvenceBackbone=Mistral-7B2025.03 | 92 | — | |
| OSCAR-llamaBackbone=Mistral-7B2025.03 | 92 | — | |
| OSCAR-8-LayersBackbone=Mistral-7B2025.03 | 92 | — | |
| No compressionBackbone=Mistral-24B2025.03 | 92 | — | |
| GPT-3.5-turbo-1106Retrieval-Augmented Generation=Disabled2024.07 | 91.7 | 82.9 | |
| Llama3-ChatQA-1.5 70BRetrieval-Augmented Generation=Enabled2024.07 | 91.4 | 85.6 | |
| GPT-4-turbo-2024-0409 RAGRetrieval-Augmented Generation=Enabled2024.07 | 91.1 | 70.2 | |
| OSCAR-5-LayersBackbone=Mistral-7B2025.03 | 91 | — | |
| OSCAR-8-LayersBackbone=Qwen-7B2025.03 | 91 | — | |
| OSCAR-llamaBackbone=Qwen-7B2025.03 | 91 | — | |
| PISCOBackbone=Mistral-7B2025.03 | 90 | — | |
| No compressionBackbone=Qwen-7B2025.03 | 90 | — | |
| Llama3-RankRAG 8BRetrieval-Augmented Generation=Enabled, Zero-shot=true2024.07 | 89.5 | 82.9 | |
| Llama3-Instruct 70BRetrieval-Augmented Generation=Enabled2024.07 | 89.3 | 82.4 | |
| GPT-4-0613 RAGRetrieval-Augmented Generation=Enabled2024.07 | 88.5 | 75 | |
| GPT-3.5-turbo-1106 RAGRetrieval-Augmented Generation=Enabled2024.07 | 88 | 79.7 | |
| Llama3-ChatQA-1.5 8BRetrieval-Augmented Generation=Enabled2024.07 | 87.6 | 81 | |
| OSCAR-5-LayersBackbone=Llama-1B2025.03 | 86 | — | |
| No compressionBackbone=Llama-1B2025.03 | 82 | — | |
| Llama3-Instruct 8BRetrieval-Augmented Generation=Enabled2024.07 | 80.4 | 70.7 | |
| No RAGBackbone=Mistral-7B2025.03 | 79 | — | |
| Gemma 3 12Bshots=5-shot2026.01 | 78.8 | — | |
| LatentMemMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 74.92 | — | |
| Ministral 3 14Bshots=5-shot2026.01 | 74.9 | — | |
| G-MemoryMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 74.6 | — | |
| VoyagerMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 74.53 | — | |
| LatentMemMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 74.45 | — | |
| No SteeringBase Model=Llama 3 8B, Method Category=Baselines, Random Seeds=52025.12 | 74.2 | — | |
| MetaGPTMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 74.1 | — | |
| LatentMemMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 74.1 | — | |
| LatentMemMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 74 | — | |
| G-MemoryMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 74 | — | |
| MetaGPTMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 73.87 | — | |
| GenerativeMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 73.77 | — | |
| G-MemoryMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 73.69 | — | |
| VoyagerMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 73.69 | — | |
| OAgentMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 73.52 | — | |
| OAgentMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 73.49 | — | |
| GenerativeMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 73.22 | — | |
| InstructGPTRetrieval-Augmented Generation=Disabled2024.07 | 73.2 | 65.8 | |
| OAgentMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 73.12 | — | |
| MetaGPTMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 72.99 | — | |
| No-memoryMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 72.86 | — | |
| GenerativeMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 72.83 | — | |
| No-memoryMAS Framework=DyLAN, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 72.62 | — | |
| MetaGPTMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 72.57 | — | |
| VoyagerMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 72.57 | — | |
| Prompt guardrailsBase Model=Llama 3 8B, Method Category=Baselines, Random Seeds=52025.12 | 72.4 | — | |
| GenerativeMAS Framework=CAMEL, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-out2026.02 | 72.36 | — | |
| OAgentMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 72.2 | — | |
| VoyagerMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 72.17 | — | |
| No-memoryMAS Framework=AutoGen, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 72.03 | — | |
| Gradient CuffBase Model=Llama 3 8B, Method Category=Baselines, Random Seeds=52025.12 | 71.8 | — | |
| G-MemoryMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 71.79 | — | |
| No-memoryMAS Framework=MacNet, LLM Backbone=Llama-3.1-8B-Instruct, MAS Framework Evaluation Setting=Held-in2026.02 | 71.62 | — | |
| MixtralParameters=13B/47B, Complexity Type=Quadratic, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 71 | — | |
| Qwen 3 14Bshots=5-shot2026.01 | 70.3 | — | |
| GSAEBase Model=Llama 3 8B, Method Category=Main Method, Random Seeds=52025.12 | 70 | — | |
| Self-RAG 13BRetrieval-Augmented Generation=Enabled2024.07 | 69.3 | — | |
| Input Gate OnlyBase Model=Llama 3 8B, Method Category=Ablation, Random Seeds=52025.12 | 68.5 | — | |
| Ministral 3 8Bshots=5-shot2026.01 | 68.1 | — | |
| MoST2026.01 | 67.1 | — | |
| Self-RAG 7BRetrieval-Augmented Generation=Enabled2024.07 | 66.4 | — | |
| MinMo2026.01 | 65.8 | — | |
| Llama3Parameters=8B, Complexity Type=Quadratic, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 65.78 | — | |
| GSAE-1DBase Model=Llama 3 8B, Method Category=Ablation, Random Seeds=52025.12 | 65.3 | — | |
| Gemma 3 4Bshots=5-shot2026.01 | 64 | — | |
| Qwen 3 8Bshots=5-shot2026.01 | 63.9 | — | |
| No gatingBase Model=Llama 3 8B, Method Category=Ablation, Random Seeds=52025.12 | 63.2 | — | |
| SAE steeringBase Model=Llama 3 8B, Method Category=Baselines, Random Seeds=52025.12 | 62.2 | — | |
| SafeSwitchBase Model=Llama 3 8B, Method Category=Baselines, Random Seeds=52025.12 | 61 | — | |
| DIFFUSPEECHType=Diff., Evaluation Mode=text-in, text-out2026.01 | 60.3 | — | |
| CAABase Model=Llama 3 8B, Method Category=Baselines, Random Seeds=52025.12 | 60.1 | — | |
| Ministral 3 3Bshots=5-shot2026.01 | 59.2 | — | |
| Phi-4-MultimodalType=AR, Evaluation Mode=text-in, text-out2026.01 | 58.5 | — | |
| SpikingBrain-7BParameters=7B, Complexity Type=Linear, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 57.03 | — | |
| Ours (SCD)decoding=SCD2025.02 | 56.3 | — | |
| Greedydecoding=Greedy2025.02 | 56 | — | |
| Qwen2.5Parameters=7B, Complexity Type=Quadratic, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 55.72 | — | |
| LLaDAType=Diff., Evaluation Mode=text-in, text-out2026.01 | 55.6 | — | |
| SpikingBrain-76BParameters=12B/76B, Complexity Type=Hybrid, Evaluation Framework=vLLM, Evaluation Method=generation-based2025.09 | 55.13 | — | |
| MinMoType=AR, Evaluation Mode=text-in, text-out2026.01 | 54.8 | — | |
| Qwen 3 4Bshots=5-shot2026.01 | 53 | — | |
| MOUE 16A2Act.=2.2B, VP.=14.6B, Adaptation Protocol=supervised fine-tuning, Base Model=JetMoE-8E2026.03 | 51.2 | — | |
| MOUE 10A2Act.=2.2B, VP.=9.7B, Adaptation Protocol=supervised fine-tuning, Base Model=JetMoE-8E2026.03 | 50.7 | — | |
| MOUE 12A2Act.=2.2B, VP.=11.3B, Adaptation Protocol=supervised fine-tuning, Base Model=JetMoE-8E2026.03 | 50.5 | — | |
| Vanilla MoEAct.=2.2B, VP.=8.0B, Adaptation Protocol=supervised fine-tuning, Base Model=JetMoE-8E2026.03 | 49.5 | — | |
| MoshiType=AR, Evaluation Mode=text-in, text-out2026.01 | 48.5 | — | |
| Moshi2026.01 | 48.5 | — | |
| MOUE 76A8Act.=1.3B, VP.=8.1B, Adaptation Protocol=supervised fine-tuning, Base Model=OLMoE-64E2026.03 | 47.6 | — | |
| Vanilla MoEAct.=1.3B, VP.=6.9B, Adaptation Protocol=supervised fine-tuning, Base Model=OLMoE-64E2026.03 | 45.3 | — | |
| MOUE 72A8Act.=1.3B, VP.=7.7B, Adaptation Protocol=supervised fine-tuning, Base Model=OLMoE-64E2026.03 | 44.5 | — | |
| MOUE 68A8Act.=1.3B, VP.=7.3B, Adaptation Protocol=supervised fine-tuning, Base Model=OLMoE-64E2026.03 | 44.1 | — | |
| SpiritLMType=AR, Evaluation Mode=text-in, text-out2026.01 | 42 | — |