Multiple Choice Question Answering on MMLU (English/Spanish Gap Analysis)
59.7EN AccuracyNemotron-3 120B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Nemotron-3 120BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 59.7 | 56.5 | 3.2 | 1.9 | |
| Nemotron-3 120BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 58.4 | 53.3 | 5 | — | |
| Qwen2.5 72BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 55.4 | 51.3 | 4.1 | 3.2 | |
| Qwen2.5 72BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 54.6 | 47.3 | 7.3 | — | |
| LLaMA-3.1 70BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 54.1 | 49.6 | 4.5 | 0.4 | |
| LLaMA-3.1 70BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=true2026.06 | 53.1 | 48.1 | 5 | — | |
| GPT-OSS 120BEvaluation Protocol=ParaEval (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=false2026.06 | 50 | 43.3 | 6.7 | 0.5 | |
| GPT-OSS 120BEvaluation Protocol=Original (raw LL), Shots=5-shot, Scoring=raw LL, Pretrained-only=false2026.06 | 48.6 | 41.4 | 7.2 | — | |
| ParaEvalScale=8B2026.06 | 42.1 | 41.7 | 0.4 | — | |
| Standard EvaluationScale=8B2026.06 | 41.2 | 37.8 | 3.4 | — | |
| ParaEvalScale=3B2026.06 | 38.3 | 37.8 | 0.52 | — | |
| Standard EvaluationScale=3B2026.06 | 37.3 | 33.5 | 3.8 | — | |
| ParaEvalScale=1B2026.06 | 34.41 | 33.92 | 0.49 | — | |
| Standard EvaluationScale=1B2026.06 | 33.54 | 30.87 | 2.68 | — |