Multiple Choice Question Answering on ARC Challenge (Bilingual/Gap Analysis)
67.1EN AccuracyNemotron-3 120B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Nemotron-3 120BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 67.1 | 64.3 | 2.7 | 1.3 | |
| LLaMA-3.1 70BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 66.5 | 60.2 | 6.3 | 1 | |
| Nemotron-3 120BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 66.5 | 62.5 | 4 | — | |
| Qwen2.5 72BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 66.2 | 63.1 | 3.1 | 1.4 | |
| LLaMA-3.1 70BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 66.2 | 58.9 | 7.3 | — | |
| Qwen2.5 72BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 65.8 | 61.3 | 4.4 | — | |
| GPT-OSS 120BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=false2026.06 | 57.9 | 48 | 9.9 | 2.3 | |
| GPT-OSS 120BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=false2026.06 | 57.4 | 45.2 | 12.2 | — | |
| ParaEvalScale=8B2026.06 | 56.7 | 56.2 | 0.5 | — | |
| Standard EvaluationScale=8B2026.06 | 54.4 | 53.6 | 0.8 | — | |
| ParaEvalScale=3B2026.06 | 48.3 | 47.8 | 0.51 | — | |
| Standard EvaluationScale=3B2026.06 | 46.9 | 45.1 | 1.88 | — | |
| ParaEvalScale=1B2026.06 | 39.79 | 39.76 | 0.03 | — | |
| Standard EvaluationScale=1B2026.06 | 38.62 | 37.49 | 1.14 | — |