Multitask Knowledge Evaluation on MMLU-Pro
80Pass@1Qwen3
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3Model Category=Original Models, Sampling Temperature=0.6, Shuffle options=true2026.05 | 80 | |
| Qwen3-8192Model Category=Original Models, Max Generation Length=8192, Sampling Temperature=0.6, Shuffle options=true2026.05 | 76.5 | |
| LUFFYModel Category=Off-Policy and Mixed-Policy, Sampling Temperature=0.6, Shuffle options=true2026.05 | 50.1 | |
| TGPO-annealingModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 50.1 | |
| TGPOModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 48.9 | |
| TGPORModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 48.1 | |
| GRPO++Model Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 46.9 | |
| KDRLModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 46.9 | |
| SFTModel Category=Off-Policy and Mixed-Policy, Sampling Temperature=0.6, Shuffle options=true2026.05 | 44.9 | |
| Oat-ZeroModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 41.7 | |
| SimpleRL-ZeroModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 34.5 | |
| PRIME-ZeroModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 32.7 | |
| OP DistillModel Category=On-Policy Methods, Sampling Temperature=0.6, Shuffle options=true2026.05 | 23 | |
| Qwen2.5-Math-7BModel Category=Original Models, Sampling Temperature=0.6, Shuffle options=true2026.05 | 16.9 |