Multiple Choice Question Answering on ARC Easy (EN/ES Accuracy and Gaps)
89.7EN AccuracyLLaMA-3.1 70B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaMA-3.1 70BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 89.7 | 85.8 | 3.9 | — | |
| Qwen2.5 72BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 89.4 | 86.4 | 3 | — | |
| Qwen2.5 72BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 89.1 | 87 | 2.2 | 0.8 | |
| LLaMA-3.1 70BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 88.9 | 86.2 | 2.7 | 1.2 | |
| Nemotron-3 120BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 88.3 | 88 | 0.3 | 1.4 | |
| Nemotron-3 120BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 88.2 | 86.5 | 1.7 | — | |
| GPT-OSS 120BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=false2026.06 | 84.7 | 79.7 | 5 | — | |
| GPT-OSS 120BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=false2026.06 | 84.3 | 80.6 | 3.7 | 1.4 | |
| Standard EvaluationScale=8B2026.06 | 83.33 | 83.25 | 0.08 | — | |
| ParaEvalScale=8B2026.06 | 83.21 | 83.16 | 0.04 | — | |
| Standard EvaluationScale=3B2026.06 | 79.29 | 78.96 | 0.34 | — | |
| ParaEvalScale=3B2026.06 | 78.91 | 78.87 | 0.04 | — | |
| Standard EvaluationScale=1B2026.06 | 72.26 | 69.96 | 2.3 | — | |
| ParaEvalScale=1B2026.06 | 70.52 | 70.24 | 0.28 | — |