Causal Reasoning on XCOPA
94.4AccuracyPaLM 2
Evaluation Results
| Method | Links | |
|---|---|---|
| PaLM 2Number of exemplars (k-shot)=4, Model variant=Original2023.05 | 94.4 | |
| PaLM + CoTNumber of exemplars (k-shot)=4, Chain-of-thought prompting (CoT)=true2023.05 | 89.9 | |
| BYOL-nyaParameters=12B, Pre-training Type=CPT2026.01 | 71.2 | |
| BYOL-nyaParameters=4B, Pre-training Type=CPT2026.01 | 70 | |
| Apertus-8Bmode=Zero-shot, model_type=Base2026.05 | 65.69 | |
| QuaSARBackbone=Llama-3-8B2025.02 | 65 | |
| Apertus-v1.1-4Bmode=Zero-shot, model_type=Base2026.05 | 63.82 | |
| HQ seedFiltering Strategy=HQ seed2026.04 | 62 | |
| Qwen3-4B-Basemode=Zero-shot, model_type=Base2026.05 | 61.82 | |
| ML (15%)Filtering Strategy=ML (15%)2026.04 | 61.6 | |
| BYOL-nyaParameters=1B, Pre-training Type=CPT2026.01 | 61.2 | |
| Gemma-3Parameters=12B, Pre-training Type=PT2026.01 | 60.8 | |
| MLFiltering Strategy=ML2026.04 | 60.8 | |
| ML (Q3)Filtering Strategy=ML (Q3)2026.04 | 60.2 | |
| Apertus-v1.1-1.5Bmode=Zero-shot, model_type=Base2026.05 | 59.76 | |
| MKC-eFiltering Strategy=MKC-e2026.04 | 59.4 | |
| HQFiltering Strategy=HQ2026.04 | 59.2 | |
| No filteringFiltering Strategy=No filtering2026.04 | 58.6 | |
| Qwen3-1.7B-Basemode=Zero-shot, model_type=Base2026.05 | 58.35 | |
| SmolLM-3B-Basemode=Zero-shot, model_type=Base2026.05 | 58.02 | |
| Gemma-3Parameters=4B, Pre-training Type=PT2026.01 | 57.2 | |
| CoTBackbone=Llama-3-8B2025.02 | 56.9 | |
| baselineBackbone=Llama-3-8B2025.02 | 56.4 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 56 | |
| ACROS2026.05 | 55.8 | |
| EuroLLM-1.7Bmode=Zero-shot, model_type=Base2026.05 | 55.76 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.7 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 55.6 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.5 | |
| Apertus-v1.1-0.5Bmode=Zero-shot, model_type=Base2026.05 | 55.49 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.3 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.2 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 55.1 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 55 | |
| Qwen3-0.6B-Basemode=Zero-shot, model_type=Base2026.05 | 54.96 | |
| GemmaParameters=270M2026.05 | 54.8 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 54.7 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 54.4 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 54.1 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 53.9 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 53.9 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 53.8 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 53.6 | |
| SmolLM2-1.7Bmode=Zero-shot, model_type=Base2026.05 | 53.51 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 53.4 | |
| ApertusParameters=8B, Pre-training Type=25092026.01 | 53.2 | |
| Qwen2Parameters=0.5B2026.05 | 53.2 | |
| Frozen BP2026.05 | 53 | |
| Qwen-3Parameters=8B, Pre-training Type=Base2026.01 | 52.2 | |
| Qwen-3Parameters=1.7B, Pre-training Type=Base2026.01 | 52 | |
| SmolLM2Variant=Base2026.05 | 52 | |
| Gemma-3Parameters=1B, Pre-training Type=PT2026.01 | 51.8 | |
| Llama-3.1Parameters=8B2026.01 | 51.2 | |
| Qwen-3Parameters=14B, Pre-training Type=Base2026.01 | 51 | |
| Llama-3.2Parameters=1B2026.01 | 48.6 |