Mathematical Reasoning on Average (AIME24, AIME25, AMC23, MATH500)
55.4Pass@1AER
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AERModel Backbone=Qwen3-8B-Base2025.10 | 55.4 | 76 | |
| Clip-CovModel Backbone=Qwen3-8B-Base2025.10 | 55 | 75.2 | |
| Clip-HigherModel Backbone=Qwen3-8B-Base2025.10 | 51.3 | 74.5 | |
| AERModel Backbone=Qwen3-4B-Base2025.10 | 51.1 | 72.5 | |
| Clip-CovModel Backbone=Qwen3-4B-Base2025.10 | 50.1 | 70.6 | |
| KL-CovModel Backbone=Qwen3-8B-Base2025.10 | 48.4 | 72.5 | |
| Ent-AdvModel Backbone=Qwen3-8B-Base2025.10 | 46.3 | 63.6 | |
| GRPOModel Backbone=Qwen3-8B-Base2025.10 | 46 | 66 | |
| KL-CovModel Backbone=Qwen3-4B-Base2025.10 | 45.6 | 68.6 | |
| Clip-HigherModel Backbone=Qwen3-4B-Base2025.10 | 45.3 | 63.9 | |
| GRPOModel Backbone=Qwen3-4B-Base2025.10 | 43.9 | 64 | |
| Ent-AdvModel Backbone=Qwen3-4B-Base2025.10 | 40.7 | 62.6 | |
| BaseModel Backbone=Qwen3-8B-Base2025.10 | 33.1 | 66.6 | |
| BaseModel Backbone=Qwen3-4B-Base2025.10 | 24.8 | 61.5 | |
| GRPO + LatentReviseModel=Qwen3-8B-Base, Training Condition=DAPO-14k-Hard2026.06 | 22.58 | 65.14 | |
| SFT + LatentReviseModel=Qwen3-8B-Base, Training Condition=DAPO-14k-Hard2026.06 | 21.26 | 64.46 | |
| GRPO(n=16)Model=Qwen3-8B-Base, Training Condition=DAPO-14k-Hard2026.06 | 21.01 | 58.9 | |
| GRPOModel=Qwen3-8B-Base, Training Condition=DAPO-14k-Hard2026.06 | 19.43 | 62.75 | |
| BaseModel=Qwen3-8B-Base, Training Condition=DAPO-14k-Hard2026.06 | 18.24 | 64.36 | |
| GRPO + LatentReviseModel=Qwen2.5-1.5B-Inst, Training Condition=DAPO-14k-Hard2026.06 | 18.13 | 53.51 | |
| GRPO(n=16)Model=Qwen2.5-1.5B-Inst, Training Condition=DAPO-14k-Hard2026.06 | 17.71 | 50.69 | |
| SFT + LatentReviseModel=Qwen2.5-1.5B-Inst, Training Condition=DAPO-14k-Hard2026.06 | 17.62 | 50.5 | |
| OPSDModel=Qwen3-8B-Base, Training Condition=DAPO-14k-Hard2026.06 | 16.09 | 60.36 | |
| BaseModel=Qwen2.5-1.5B-Inst, Training Condition=DAPO-14k-Hard2026.06 | 16.05 | 46.67 | |
| GRPOModel=Qwen2.5-1.5B-Inst, Training Condition=DAPO-14k-Hard2026.06 | 14.65 | 44.84 | |
| OPSDModel=Qwen2.5-1.5B-Inst, Training Condition=DAPO-14k-Hard2026.06 | 7.07 | 40.42 |