Mathematics on AIME 25 (Score (%))
61.11Score (%)Qwen3.5-9B + AR-SFT
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-9B + AR-SFTModel Size=9B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 61.11 | |
| Qwen3.5-9B (released)Model Size=9B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 60 | |
| FLARE-9BModel Size=9B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 54.44 | |
| Qwen3.5-4B (released)Model Size=4B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 48.89 | |
| Qwen3.5-4B + AR-SFTModel Size=4B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 46.67 | |
| FLARE-4BModel Size=4B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 43.33 | |
| MPO+Backbone=olmo-3.1-32b-it-sft2026.04 | 35.8 | |
| Mag+Backbone=olmo-3.1-32b-it-sft2026.04 | 33.3 | |
| Soft+Backbone=olmo-3.1-32b-it-sft2026.04 | 32.3 | |
| Pair+Backbone=olmo-3.1-32b-it-sft2026.04 | 31.7 | |
| Mag+Backbone=olmo-3-7b-it-sft2026.04 | 28.1 | |
| RFTBackbone=olmo-3.1-32b-it-sft2026.04 | 27.5 | |
| Pair+Backbone=olmo-3-7b-it-sft2026.04 | 26.9 | |
| FLARE-2BModel Size=2B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 26.67 | |
| Soft+Backbone=olmo-3-7b-it-sft2026.04 | 26.5 | |
| MPO+Backbone=olmo-3-7b-it-sft2026.04 | 26 | |
| DPOBackbone=olmo-3.1-32b-it-sft2026.04 | 25.6 | |
| DPOBackbone=olmo-3-7b-it-sft2026.04 | 25.2 | |
| Qwen3.5-2B + AR-SFTModel Size=2B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 22.22 | |
| RFTBackbone=olmo-3-7b-it-sft2026.04 | 17.3 | |
| Qwen3.5-2B (released)Model Size=2B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 12.22 | |
| BaseBackbone=olmo-3.1-32b-it-sft2026.04 | 8.1 | |
| BaseBackbone=olmo-3-7b-it-sft2026.04 | 7.9 |