Open-domain Question Answering on TriviaQA (Accuracy)
81.4AccuracyPaLM
Evaluation Results
| Method | Links | |
|---|---|---|
| PaLMModel Size (Billions of Parameters)=540, Num Tokens (Billions of Tokens)=780, Training FLOP Count (Zettaflops)=2527.2, Evaluation Protocol=few-shot2022.04 | 81.4 | |
| PaLMModel Size (Billions of Parameters)=540, Num Tokens (Billions of Tokens)=780, Training FLOP Count (Zettaflops)=2527.2, Evaluation Protocol=0-shot2022.04 | 76.9 | |
| ChinchillaModel Size (Billions of Parameters)=70, Num Tokens (Billions of Tokens)=1400, Training FLOP Count (Zettaflops)=588.0, Evaluation Protocol=few-shot2022.04 | 73.2 | |
| LLaMA 2 70BActive Params=70B2024.01 | 73 | |
| PaLMModel Size (Billions of Parameters)=62, Num Tokens (Billions of Tokens)=795, Training FLOP Count (Zettaflops)=295.7, Evaluation Protocol=few-shot2022.04 | 72.7 | |
| Mixtral 8x7BActive Params=13B2024.01 | 71.5 | |
| LLaMA 1 33BActive Params=33B2024.01 | 68.5 | |
| PaLMModel Size (Billions of Parameters)=62, Num Tokens (Billions of Tokens)=795, Training FLOP Count (Zettaflops)=295.7, Evaluation Protocol=0-shot2022.04 | 67.3 | |
| ChinchillaModel Size (Billions of Parameters)=70, Num Tokens (Billions of Tokens)=1400, Training FLOP Count (Zettaflops)=588.0, Evaluation Protocol=0-shot2022.04 | 67 | |
| LLaMA 2 13BActive Params=13B2024.01 | 64 | |
| GopherModel Size (Billions of Parameters)=280, Num Tokens (Billions of Tokens)=300, Training FLOP Count (Zettaflops)=504.0, Evaluation Protocol=few-shot2022.04 | 63.6 | |
| Mistral 7BActive Params=7B2024.01 | 62.5 | |
| MoE with MPIModel Scale=11B, Setup=5-shot2026.06 | 56.89 | |
| LLaMA 2 7BActive Params=7B2024.01 | 56.6 | |
| MoEModel Scale=11B, Setup=5-shot2026.06 | 55.41 | |
| GopherModel Size (Billions of Parameters)=280, Num Tokens (Billions of Tokens)=300, Training FLOP Count (Zettaflops)=504.0, Evaluation Protocol=0-shot2022.04 | 52.8 | |
| PaLMModel Size (Billions of Parameters)=8, Num Tokens (Billions of Tokens)=780, Training FLOP Count (Zettaflops)=37.4, Evaluation Protocol=few-shot2022.04 | 48.5 | |
| MoE with MPIModel Scale=3B, Setup=5-shot2026.06 | 46.52 | |
| MoEModel Scale=3B, Setup=5-shot2026.06 | 45.78 | |
| PaLMModel Size (Billions of Parameters)=8, Num Tokens (Billions of Tokens)=780, Training FLOP Count (Zettaflops)=37.4, Evaluation Protocol=0-shot2022.04 | 39.5 |