Language Understanding on MMLU-Pro
80.6Accuracygpt-oss-120B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| gpt-oss-120BReasoning effort=High, KV cache precision=FP82026.02 | 80.6 | — | |
| gpt-oss-120BReasoning effort=High, KV cache precision=BF162026.02 | 80.41 | — | |
| gpt-oss-puzzle-88BReasoning effort=High, KV cache precision=BF162026.02 | 79.32 | — | |
| gpt-oss-puzzle-88BReasoning effort=High, KV cache precision=FP82026.02 | 79.19 | — | |
| gpt-oss-120BReasoning effort=Medium, KV cache precision=BF162026.02 | 78.86 | — | |
| gpt-oss-120BReasoning effort=Medium, KV cache precision=FP82026.02 | 78.71 | — | |
| gpt-oss-puzzle-88BReasoning effort=Medium, KV cache precision=FP82026.02 | 78.18 | — | |
| gpt-oss-puzzle-88BReasoning effort=Medium, KV cache precision=BF162026.02 | 78.03 | — | |
| gpt-oss-puzzle-88BReasoning effort=Low, KV cache precision=FP82026.02 | 75.62 | — | |
| gpt-oss-puzzle-88BReasoning effort=Low, KV cache precision=BF162026.02 | 75.56 | — | |
| gpt-oss-120BReasoning effort=Low, KV cache precision=BF162026.02 | 75.18 | — | |
| gpt-oss-120BReasoning effort=Low, KV cache precision=FP82026.02 | 75.04 | — | |
| NBDiff-7B-INSTRUCTParameters=7B, Training Protocol=Instruct, Sampling=Standard2025.12 | 71.9 | — | |
| Llama 3 405BModel Type=Pre-trained2024.07 | 61.6 | — | |
| MOCEdge Density=0.5, Backbone=Gemma-2-27B-Instruct2026.06 | 59.29 | — | |
| NBDiff-7B-BASEConfiguration=Base2025.12 | 59.1 | — | |
| MOCEdge Density=1.0, Backbone=Gemma-2-27B-Instruct2026.06 | 58.21 | — | |
| Vanilla MASEdge Density=0.5, Backbone=Gemma-2-27B-Instruct2026.06 | 57.86 | — | |
| MOCEdge Density=0.3, Backbone=Gemma-2-27B-Instruct2026.06 | 57.5 | — | |
| SDAR 8BParameters=8B, Sampling=Standard2025.12 | 56.9 | — | |
| Vanilla MASEdge Density=0.3, Backbone=Gemma-2-27B-Instruct2026.06 | 56.79 | — | |
| Vanilla MASEdge Density=1.0, Backbone=Gemma-2-27B-Instruct2026.06 | 56.43 | — | |
| MOCEdge Density=0.7, Backbone=Gemma-2-27B-Instruct2026.06 | 56.07 | — | |
| Qwen2-72B2024.07 | 55.6 | — | |
| Vanilla MASEdge Density=0.7, Backbone=Gemma-2-27B-Instruct2026.06 | 55.36 | — | |
| OpenReasoner-ZeroBase Model=Qwen2.5-Math-7B-Base2025.05 | 54.6 | — | |
| LoopRPTParameters=2.6B2026.03 | 54.19 | — | |
| OuroParameters=2.6B2026.03 | 54.07 | — | |
| Llama 3 70BModel Type=Pre-trained2024.07 | 53.8 | — | |
| Qwen3Parameters=8B2026.03 | 53.72 | — | |
| Llama-3-70B2024.07 | 52.8 | — | |
| Mistral 2501 BaseFew-shot=5 shot2026.02 | 52 | — | |
| Mixtral 8x22BModel Type=Pre-trained2024.07 | 51.5 | — | |
| Qwen2.5Parameters=4B2026.03 | 51.4 | — | |
| Mixtral-8x22B2024.07 | 49.5 | — | |
| TemplateRLBase Model=Qwen2.5-Math-7B-Base2025.05 | 49.5 | — | |
| Qwen1.5-110B2024.07 | 49.4 | — | |
| Gemma3Parameters=12B2026.03 | 49.21 | — | |
| LLaDA2.0-mini preview 16BA1BParameters=16BA1B, Sampling=Standard2025.12 | 49.2 | — | |
| FiMi BaseFew-shot=5 shot2026.02 | 49 | — | |
| Yi-1.5-34BArchitecture=Dense, # Act Params=32B, # Params=32B2024.07 | 48.3 | — | |
| Dream-v0Configuration=Base-7B2025.12 | 48.2 | — | |
| Qwen1.5-72B2024.07 | 45.8 | — | |
| Single LLMBackbone=Gemma-2-27B-Instruct2026.06 | 45.71 | — | |
| LLaDA-MoE 7B-A1BParameters=7B-A1B, Sampling=Standard2025.12 | 44.6 | — | |
| Qwen1.5-32BArchitecture=Dense, # Act Params=34B, # Params=34B2024.07 | 44 | — | |
| Qwen2.5Parameters=7B2026.03 | 43.55 | — | |
| Dream-v0 Instruct-7BParameters=7B, Training Protocol=Instruct, Sampling=Standard2025.12 | 43.3 | — | |
| Llama3.1Parameters=8B2026.03 | 43.24 | — | |
| Qwen2-57B-A14BArchitecture=MoE, # Act Params=14B, # Params=57B2024.07 | 43 | — | |
| GRPOBase Model=Qwen2.5-Math-7B-Base2025.05 | 42.1 | — | |
| LLaDA-8BConfiguration=Base2025.12 | 41.8 | — | |
| Oat-ZeroBase Model=Qwen2.5-Math-7B-Base2025.05 | 41.8 | — | |
| Llama3-8Bfine-tuning=LoRA2026.02 | 41.67 | — | |
| Krause-Llama3-8B2026.02 | 41.67 | — | |
| S-SimPOIteration=Iter12025.12 | 41.41 | — | |
| Mixtral-8x7BArchitecture=MoE, # Act Params=12B, # Params=47B2024.07 | 41 | — | |
| S-SimPOIteration=Iter02025.12 | 40.49 | — | |
| LLaDA-MoE-7BConfiguration=A1B-Base2025.12 | 39.2 | — | |
| Gemma3Parameters=3B2026.03 | 37.87 | — | |
| Krause-Qwen1.5-7BWindow size=32, top-k=162026.02 | 37.5 | — | |
| Krause-Qwen1.5-7BWindow size=18, top-k=122026.02 | 37.5 | — | |
| Llama3-8B2026.02 | 37.5 | — | |
| Llama 3 8BModel Type=Pre-trained2024.07 | 37.1 | — | |
| SDAR-1.7BScale=1.7B, Tokens=55B, FLOPs (1e18)=561.02026.06 | 37 | — | |
| SimpleRL-ZeroBase Model=Qwen2.5-Math-7B-Base2025.05 | 36.9 | — | |
| Qwen1.5-7BFine-tuning protocol=LoRA2026.02 | 36.11 | — | |
| GRPOModel=Llama3.2-3B-Ins2025.05 | 36.1 | — | |
| Qwen2.5-Math-7B-Instruct2025.05 | 36.1 | — | |
| INTUITORModel=Llama3.2-3B-Ins2025.05 | 35.8 | — | |
| GRPO-PVModel=Llama3.2-3B-Ins2025.05 | 35.2 | — | |
| Gemma 7BModel Type=Pre-trained2024.07 | 35.1 | — | |
| Krause-Qwen1.5-7BWindow size=48, top-k=242026.02 | 34.72 | — | |
| Qwen3Parameters=4B2026.03 | 34.61 | — | |
| PRIME-ZeroBase Model=Qwen2.5-Math-7B-Base2025.05 | 34.3 | — | |
| CoPE2026.02 | 34.05 | — | |
| BaselineModel=Llama3.2-3B-Ins2025.05 | 34 | — | |
| HardClip2026.02 | 33.95 | — | |
| RoPE2026.02 | 33.52 | — | |
| Llama3.2Parameters=3B2026.03 | 33.34 | — | |
| Mistral 7BModel Type=Pre-trained2024.07 | 32.5 | — | |
| OPDLM-1.7BScale=1.7B, Tokens=0.072B, FLOPs (1e18)=0.982026.06 | 31.1 | — | |
| Qwen1.5-7BFine-tuning protocol=None2026.02 | 30.56 | — | |
| GRPOModel=OLMo2-7B-SFT2025.05 | 29.6 | — | |
| BaselineModel=OLMo2-7B-SFT2025.05 | 29.5 | — | |
| S-IPOIteration=Iter42025.12 | 29.39 | — | |
| S-IPOIteration=Iter32025.12 | 29.29 | — | |
| S-IPOIteration=Iter02025.12 | 29.22 | — | |
| S-IPOIteration=Iter22025.12 | 29.11 | — | |
| INTUITORModel=OLMo2-7B-SFT2025.05 | 29.1 | — | |
| S-IPOIteration=Iter12025.12 | 29.01 | — | |
| SPACEIteration=Iter02025.12 | 28.81 | — | |
| S-SimPOIteration=Iter42025.12 | 28.77 | — | |
| SPACEIteration=Iter32025.12 | 28.76 | — | |
| SPACEIteration=Iter22025.12 | 28.72 | — | |
| SPACEIteration=Iter42025.12 | 28.72 | — | |
| SPINIteration=Iter22025.12 | 28.67 | — | |
| S-SimPOIteration=Iter22025.12 | 28.65 | — | |
| S-SimPOIteration=Iter32025.12 | 28.62 | — | |
| SPINIteration=Iter32025.12 | 28.57 | — |