Reasoning on SIQA
83.2AccuracySupervised SOTA
Evaluation Results
| Method | Links | |
|---|---|---|
| Supervised SOTAsupervised=true2022.03 | 83.2 | |
| DRAGON2022.10 | 76.8 | |
| RoBERTa2022.10 | 75.9 | |
| QAGNN2022.10 | 75.7 | |
| GreaseLM2022.10 | 75.5 | |
| phi-1.5-web (1.3B)Zero-shot=true, Number of Parameters=1.3B2023.09 | 53 | |
| phi-1.5 (1.3B)Zero-shot=true, Number of Parameters=1.3B2023.09 | 52.6 | |
| LLaMAParameters=65B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 52.3 | |
| Chinchillazero-shot=true2022.03 | 51.3 | |
| ChinchillaParameters=70B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 51.3 | |
| Gopherzero-shot=true2022.03 | 50.6 | |
| GopherParameters=280B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 50.6 | |
| LLaMAParameters=13B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 50.4 | |
| LLaMAParameters=33B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 50.4 | |
| LLaMAParameters=7B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 48.9 | |
| Llama2-7BZero-shot=true, Number of Parameters=7B2023.09 | 48 | |
| Self-Improving PretrainingTraining Dataset=SlimPajama, Pretraining Objective=Factuality2026.01 | 46.8 | |
| Llama-7BZero-shot=true, Number of Parameters=7B2023.09 | 46.6 | |
| Self-Improving PretrainingTraining Dataset=SlimPajama, Pretraining Objective=Quality2026.01 | 46.1 | |
| Falcon-7BZero-shot=true, Number of Parameters=7B2023.09 | 45.2 | |
| MPT-7BZero-shot=true, Number of Parameters=7B2023.09 | 45.1 | |
| Qwen 2.5Params.=1.5B2025.06 | 44.9 | |
| Qwen 1.5Params.=1.8B, Tokens=2.4T2025.06 | 44.5 | |
| Muon (OSP)Params.=1.4B, Tokens=1T2025.06 | 44.4 | |
| Qwen 2Params.=1.5B, Tokens=7T2025.06 | 44.2 | |
| OLMoParams.=1.2B, Tokens=3T2025.06 | 44.1 | |
| SmolLMParams.=1.7B, Tokens=1T2025.06 | 44.1 | |
| Self-Improving PretrainingTraining Dataset=RedPajama, Pretraining Objective=Safety2026.01 | 44.1 | |
| Vicuna-13B (v1.1)Zero-shot=true, Number of Parameters=13B, Version=v1.12023.09 | 43.7 | |
| AdamParams.=1.4B, Tokens=1T2025.06 | 43.6 | |
| PythiaParams.=1.4B, Tokens=0.3T2025.06 | 43.5 | |
| LLAMA 3.2Params.=1.2B2025.06 | 43.5 | |
| Stable LM 2Params.=1.6B, Tokens=2T2025.06 | 43.5 | |
| SmolLM 2Params.=1.7B, Tokens=11T2025.06 | 43.4 | |
| MobileLLAMAParams.=1.4B, Tokens=1.3T2025.06 | 43 | |
| OPTParams.=1.3B, Tokens=0.3T2025.06 | 42.3 | |
| Llama Pretrain BaselineTraining Dataset=SlimPajama, Pretraining Objective=Standard next token prediction2026.01 | 42.2 | |
| Llama Pretrain BaselineTraining Dataset=RedPajama, Pretraining Objective=Safety2026.01 | 41.5 | |
| phi-1.5-web-only (1.3B)Zero-shot=true, Number of Parameters=1.3B2023.09 | 41.4 | |
| TinyLlamaParams.=1.1B, Tokens=2T2025.06 | 41.3 | |
| Llama BaseTraining Dataset=Original Llama Pretraining, Pretraining Objective=Standard next token prediction2026.01 | 41 | |
| Falcon-rw-1.3BZero-shot=true, Number of Parameters=1.3B2023.09 | 40.5 | |
| GPT-Neo-2.7BZero-shot=true, Number of Parameters=2.7B2023.09 | 40 | |
| GPT2-XL-1.5BZero-shot=true, Number of Parameters=1.5B2023.09 | 39.4 |