Mathematical Reasoning on Aggregate (MATH500, AMC, OlympiadBench, AIME24, AIME25)
66.6Average ScoreSWA8k-RL-1200
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SWA8k-RL-1200Training Stage=RL, RL Steps=1200, RL Training Budget=~500 GPU Hours, Attention Architecture=Sliding Window Attention, Sliding Window Size=8k, T_Train=466, T_Eval=2.232026.06 | 66.6 | 0.7 | |
| SWA4k-RL-1400Training Stage=RL, RL Steps=1400, RL Training Budget=~500 GPU Hours, Attention Architecture=Sliding Window Attention, Sliding Window Size=4k, T_Train=474, T_Eval=1.872026.06 | 66 | 0.1 | |
| SA-RL-900Training Stage=RL, RL Steps=900, Attention Architecture=Standard Attention, T_Train=498, T_Eval=2.812026.06 | 65.9 | — | |
| SWA8k-RL-900Training Stage=RL, RL Steps=900, Attention Architecture=Sliding Window Attention, Sliding Window Size=8k, T_Train=337, T_Eval=2.112026.06 | 65.5 | -0.4 | |
| SWA4k-RL-900Training Stage=RL, RL Steps=900, Attention Architecture=Sliding Window Attention, Sliding Window Size=4k, T_Train=285, T_Eval=1.842026.06 | 63.5 | -2.4 | |
| SWA2k-RL-1700Training Stage=RL, RL Steps=1700, RL Training Budget=~500 GPU Hours, Attention Architecture=Sliding Window Attention, Sliding Window Size=2k, T_Train=470, T_Eval=1.332026.06 | 63.4 | -2.5 | |
| SWA2k-RL-900Training Stage=RL, RL Steps=900, Attention Architecture=Sliding Window Attention, Sliding Window Size=2k, T_Train=225, T_Eval=1.462026.06 | 59.6 | -6.3 | |
| SA-SFTTraining Stage=SFT, Attention Architecture=Standard Attention, T_Train=615, T_Eval=10.442026.06 | 48.6 | — | |
| SWA8k-SFTTraining Stage=SFT, Attention Architecture=Sliding Window Attention, Sliding Window Size=8k, T_Train=560, T_Eval=5.542026.06 | 42.5 | -6.1 | |
| SWA4k-SFTTraining Stage=SFT, Attention Architecture=Sliding Window Attention, Sliding Window Size=4k, T_Train=528, T_Eval=4.702026.06 | 39.6 | -9 | |
| SWA2k-SFTTraining Stage=SFT, Attention Architecture=Sliding Window Attention, Sliding Window Size=2k, T_Train=507, T_Eval=4.202026.06 | 30.8 | -17.8 |