Mathematical Reasoning on AIME 2025 (avg@32)
68.44Avg@32Qwen3-Next-80B-A3B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-Next-80B-A3B-InstructArchitecture=MoE, # Total Params=80B, # Activated Params=3B2026.01 | 68.44 | |
| LongCat-Flash-LiteArchitecture=MoE + NE, # Total Params=68.5B, # Activated Params=2.9B~4.5B2026.01 | 63.23 | |
| Kimi-Linear-48B-A3BArchitecture=MoE, # Total Params=48B, # Activated Params=3B2026.01 | 59.58 | |
| Gemini 2.5 Flash-Lite2026.01 | 50.1 | |
| PMD-MEANBackbone=Qwen3-30B-A3B-Base, tau=0.1, Staleness=162026.02 | 37.19 | |
| PMD-MEANBackbone=Qwen3-30B-A3B-Base, tau=0.01, Staleness=162026.02 | 35.1 | |
| GRPOBackbone=Qwen3-30B-A3B-Base, tau=null, Staleness=162026.02 | 27.92 | |
| PMD-MEANBackbone=Qwen2.5-7B, tau=0.005, Staleness=162026.02 | 19.48 | |
| On-policyBackbone=Qwen2.5-7B, tau=null, Staleness=12026.02 | 18.33 | |
| PMD-MEANBackbone=Qwen2.5-7B, tau=0.01, Staleness=162026.02 | 17.5 | |
| TRAPOTraining Paradigm=Semi-supervised, Labeled Samples Count=4K, Unlabeled Samples Count=12K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 17.1 | |
| PMD-MEANBackbone=Qwen2.5-7B, tau=0.02, Staleness=162026.02 | 16.67 | |
| Fully SupervisedTraining Paradigm=Supervised, Labeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 15.3 | |
| Fully SupervisedTraining Paradigm=Supervised, Labeled Samples Count=4K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 14.8 | |
| TRAPOTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 13.8 | |
| Fully SupervisedTraining Paradigm=Supervised, Labeled Samples Count=1K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 13.5 | |
| TTRLTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 12.7 | |
| Token-level EntropyTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 11.9 | |
| Sentence-level EntropyTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 11.5 | |
| Self-certaintyTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 11.4 | |
| Sentence-level EntropyTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 10.7 | |
| TTRLTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 10.7 | |
| GRPOBackbone=Qwen2.5-7B, tau=null, Staleness=162026.02 | 10.52 | |
| Qwen-InstructTraining Paradigm=Original, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 10.2 | |
| Self-certaintyTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 10.2 | |
| Token-level EntropyTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 9.9 | |
| Qwen-BaseTraining Paradigm=Original, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 4.9 |