Mathematical Reasoning on AIME26 (pass@1)
40Pass@1Qwen3-4B +CAST
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-4B +CASTModel Scale=4B, Method=+CAST2026.05 | 40 | |
| Baseline ModelRollout Length=16K tokens2026.05 | 39.6 | |
| Qwen3-8B +CASTModel Scale=8B, Method=+CAST2026.05 | 36.7 | |
| Qwen3-4B BaseModel Scale=4B, Method=Base2026.05 | 23.3 | |
| Qwen3-4B +GRPOModel Scale=4B, Method=+GRPO2026.05 | 23.3 | |
| Qwen3-4B +RLSDModel Scale=4B, Method=+RLSD2026.05 | 23.3 | |
| Qwen3-4B +RLRTModel Scale=4B, Method=+RLRT2026.05 | 20 | |
| Qwen3-8B +RLSDModel Scale=8B, Method=+RLSD2026.05 | 20 | |
| FP8 Rollout P3ORollout Length=16K tokens, Training Iteration=15, Rollout Precision=FP8, Training Precision=BF162026.05 | 17.9 | |
| FP8 Rollout P3ORollout Length=16K tokens, Training Iteration=30, Rollout Precision=FP8, Training Precision=BF162026.05 | 17.5 | |
| Qwen3-4B +GRPO+OPSDModel Scale=4B, Method=+GRPO+OPSD2026.05 | 16.7 | |
| Qwen3-8B +GRPOModel Scale=8B, Method=+GRPO2026.05 | 16.7 | |
| P3ORollout Length=4K tokens, Clipping Parameter (epsilon)=∈ {0.2, 0.4, 0.6}2026.05 | 16 | |
| FP8 Rollout GRPORollout Length=16K tokens, Training Iteration=15, Rollout Precision=FP8, Training Precision=BF162026.05 | 16 | |
| Qwen3-8B BaseModel Scale=8B, Method=Base2026.05 | 13.3 | |
| Qwen3-8B +OPSDModel Scale=8B, Method=+OPSD2026.05 | 13.3 | |
| Qwen3-8B +GRPO+OPSDModel Scale=8B, Method=+GRPO+OPSD2026.05 | 13.3 | |
| Qwen3-8B +RLRTModel Scale=8B, Method=+RLRT2026.05 | 13.3 | |
| GRPO (clip avg)Rollout Length=4K tokens, Clipping Parameter (epsilon)=∈ {0.2, 0.4, 0.6}2026.05 | 12.6 | |
| Qwen3-1.7B +GRPO+OPSDModel Scale=1.7B, Method=+GRPO+OPSD2026.05 | 10 | |
| Qwen3-1.7B +CASTModel Scale=1.7B, Method=+CAST2026.05 | 10 | |
| Qwen3-4B +OPSDModel Scale=4B, Method=+OPSD2026.05 | 10 | |
| Qwen3-1.7B +GRPOModel Scale=1.7B, Method=+GRPO2026.05 | 6.7 | |
| Qwen3-1.7B +RLSDModel Scale=1.7B, Method=+RLSD2026.05 | 6.7 | |
| Qwen3-1.7B +OPSDModel Scale=1.7B, Method=+OPSD2026.05 | 3.3 | |
| Qwen3-1.7B +RLRTModel Scale=1.7B, Method=+RLRT2026.05 | 3.3 | |
| Baseline ModelRollout Length=4K tokens2026.05 | 0.6 | |
| FP8 Rollout GRPORollout Length=16K tokens, Training Iteration=30, Rollout Precision=FP8, Training Precision=BF162026.05 | 0.2 | |
| Qwen3-1.7B BaseModel Scale=1.7B, Method=Base2026.05 | 0 |