Scientific Reasoning on MMLU Pro (test)
73.21Overall AccuracyNEMOTRON-NANO-12B-V2 (Base + Post)
Evaluation Results
| Method | Links | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| NEMOTRON-NANO-12B-V2 (Base + Post)Pre-training objective=Next-token prediction, Post-training pipeline=SFT + RLVR, Evaluation protocol=Pass@1 (4 runs)2025.09 | 73.21 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NEMOTRON-NANO-12B-V2 (RLP + Post)Pre-training objective=RLP, Post-training pipeline=SFT + RLVR, Evaluation protocol=Greedy2025.09 | 67.38 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NEMOTRON-NANO-12B-V2 (RLP + Post)Pre-training objective=RLP, Post-training pipeline=SFT + RLVR, Evaluation protocol=Pass@1 (4 runs)2025.09 | 66.96 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sci-CoEModel Architecture=Qwen3-8B, Training Stage=Stage 2, Training Data Scale=30k2026.02 | 64.34 | 80.2 | 70.72 | 68.2 | 68.05 | 73.93 | 54.59 | 63.33 | 54.07 | 33.42 | 79.79 | 53.51 | 68.36 | 70.3 | 56.06 | |
| Sci-CoEModel Architecture=Qwen3-8B, Training Stage=Stage 2, Training Data Scale=18k2026.02 | 63.56 | 79.22 | 70.85 | 68.02 | 66.34 | 72.87 | 53.35 | 62.71 | 51.44 | 32.52 | 79.42 | 52.91 | 67.74 | 69.17 | 55.19 | |
| Sci-CoEModel Architecture=Qwen3-8B, Training Stage=Stage 12026.02 | 63.27 | 78.94 | 69.2 | 66.78 | 65.85 | 72.63 | 53.77 | 63.08 | 50.92 | 32.61 | 78.09 | 52.51 | 68.44 | 69.42 | 55.41 | |
| Qwen3-8BModel Architecture=Qwen3-8B, Training Stage=Base Model2026.02 | 63.19 | 78.8 | 69.71 | 68.02 | 66.1 | 72.27 | 53.04 | 62.47 | 51.97 | 31.52 | 78.53 | 51.5 | 67.67 | 69.3 | 55.95 | |
| NEMOTRON-NANO-12B-V2 (Base + Post)Pre-training objective=Next-token prediction, Post-training pipeline=SFT + RLVR, Evaluation protocol=Greedy2025.09 | 61.78 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Sci-CoEModel Architecture=Qwen2.5-7B, Training Stage=Stage 2, Training Data Scale=30k2026.02 | 58.51 | 73.92 | 68.19 | 55.39 | 61.71 | 70.62 | 42.31 | 58.19 | 48.82 | 34.06 | 72.76 | 50.5 | 59.35 | 67.17 | 54.87 | |
| Sci-CoEModel Architecture=Qwen2.5-7B, Training Stage=Stage 2, Training Data Scale=18k2026.02 | 58.05 | 73.5 | 68.69 | 57.16 | 61.22 | 68.01 | 40.04 | 58.07 | 49.08 | 32.61 | 72.54 | 50.5 | 58.97 | 66.67 | 54.65 | |
| Sci-CoEModel Architecture=Qwen2.5-7B, Training Stage=Stage 12026.02 | 57.68 | 74.34 | 67.17 | 56.89 | 60.73 | 68.72 | 40.04 | 56.6 | 47.77 | 33.33 | 72.09 | 48.5 | 59.05 | 65.54 | 53.9 | |
| Qwen2.5-7B-InstructModel Architecture=Qwen2.5-7B, Training Stage=Base Model2026.02 | 57.39 | 72.11 | 64.89 | 57.16 | 60.49 | 68.84 | 39.94 | 56.85 | 48.29 | 32.52 | 71.87 | 49.9 | 58.35 | 65.91 | 54.33 | |
| NEMOTRON-NANO-12B-V2 (RLP)Pre-training objective=RLP, Post-training pipeline=None, Evaluation protocol=Pass@1 (4 runs)2025.09 | 55.76 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NEMOTRON-NANO-12B-V2 (RLP)Pre-training objective=RLP, Post-training pipeline=None, Evaluation protocol=Greedy2025.09 | 53.13 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mistral-Small-InstructModel Architecture=Mistral-Small2026.02 | 48.4 | 71.69 | 52.72 | 36.84 | 53.66 | 60.07 | 30.55 | 53.79 | 50.13 | 31.97 | 50.85 | 48.1 | 40.34 | 63.91 | 55.09 | |
| Yi-1.5-9B-ChatModel Architecture=Yi-1.5-9B2026.02 | 45.95 | 66.67 | 54.25 | 39.49 | 50 | 60.19 | 33.23 | 43.52 | 40.94 | 26.61 | 52.48 | 40.08 | 41.42 | 59.4 | 44.91 | |
| Llama-3.1-8B-InstructModel Architecture=Llama-3.1-8B2026.02 | 44.25 | 63.04 | 49.3 | 37.63 | 48.29 | 55.09 | 29.72 | 50.73 | 42.26 | 27.25 | 43.82 | 44.49 | 40.26 | 60.03 | 44.81 | |
| Mathstral-7B-v0.1Model Architecture=Mathstral-7B2026.02 | 42 | 63.46 | 42.08 | 38.78 | 45.61 | 51.78 | 38.39 | 39.24 | 35.96 | 22.43 | 47.67 | 38.28 | 38.03 | 52.63 | 40.91 | |
| Ministral-8B-InstructModel Architecture=Ministral-8B2026.02 | 37.93 | 59 | 39.42 | 26.41 | 44.88 | 49.29 | 23.12 | 43.28 | 43.83 | 25.98 | 41.15 | 36.47 | 29.18 | 51.63 | 40.15 | |
| NEMOTRON-NANO-12B-V2 (Base)Pre-training objective=Next-token prediction, Post-training pipeline=None, Evaluation protocol=Pass@1 (4 runs)2025.09 | 27.13 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NEMOTRON-NANO-12B-V2 (Base)Pre-training objective=Next-token prediction, Post-training pipeline=None, Evaluation protocol=Greedy2025.09 | 24.16 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |