Open-domain Question Answering on Natural Questions (Accuracy)
30.6AccuracyMixtral 8x7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Mixtral 8x7BActive Params=13B2024.01 | 30.6 | |
| MistralSize=7B, Tokens=?, Shot(s)=5, Training Strategy=trained from scratch, Attention Type=softmax-attention2024.09 | 29.7 | |
| GSASize=7B, Tokens=+100B, Shot(s)=5, Training Strategy=finetuned from Mistral 7B2024.09 | 26.9 | |
| Llama2Size=7B, Tokens=2T, Shot(s)=5, Training Strategy=trained from scratch, Attention Type=softmax-attention2024.09 | 26 | |
| LLaMA 2 70BActive Params=70B2024.01 | 25.4 | |
| MambaSize=7B, Tokens=1.2T, Shot(s)=5, Training Strategy=trained from scratch2024.09 | 25.4 | |
| MoE with MPIModel Scale=11B, Setup=5-shot2026.06 | 25.36 | |
| MoEModel Scale=11B, Setup=5-shot2026.06 | 25.3 | |
| SUPRASize=7B, Tokens=+100B, Shot(s)=5, Training Strategy=finetuned from Mistral 7B2024.09 | 24.7 | |
| GemmaSize=7B, Tokens=6T, Shot(s)=5, Training Strategy=trained from scratch, Attention Type=softmax-attention2024.09 | 24.3 | |
| LLaMA 1 33BActive Params=33B2024.01 | 24.1 | |
| GSASize=7B, Tokens=+20B, Shot(s)=5, Training Strategy=finetuned from Mistral 7B2024.09 | 23.4 | |
| Mistral 7BActive Params=7B2024.01 | 23.2 | |
| GLASize=7B, Tokens=+20B, Shot(s)=5, Training Strategy=finetuned from Mistral 7B2024.09 | 22.2 | |
| RWKV6Size=7B, Tokens=1.4T, Shot(s)=5, Training Strategy=trained from scratch2024.09 | 20.9 | |
| MoE with MPIModel Scale=3B, Setup=5-shot2026.06 | 20.13 | |
| MoEModel Scale=3B, Setup=5-shot2026.06 | 17.87 | |
| LLaMA 2 7BActive Params=7B2024.01 | 17.5 | |
| LLaMA 2 13BActive Params=13B2024.01 | 16.7 | |
| RetNetSize=7B, Tokens=+20B, Shot(s)=5, Training Strategy=finetuned from Mistral 7B2024.09 | 16.2 | |
| BaseModel=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 15.12 | |
| POPModel=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | 14.65 | |
| BaseModel=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | 14.36 | |
| POPModel=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 14.01 | |
| Train on DModel=Qwen-2.5-7B, Evaluation Protocol=0-shot2026.04 | 12.54 | |
| Train on DModel=Qwen-2.5-7B-Inst, Evaluation Protocol=0-shot2026.04 | 8.5 |