Mathematical Reasoning on MATH 500 (Pass@1, Pass@8)
96Pass@1LuckyStar 111B 4-bit
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LuckyStar 111B 4-bitQuantization=4-bit, Context Window=32k2026.06 | 96 | — | |
| LuckyStar 111BContext Window=32k2026.06 | 94 | — | |
| Claude 3.7 SonnetContext Window=32k2026.06 | 92.5 | — | |
| Qwen3 235B A22BContext Window=32k2026.06 | 91.2 | — | |
| Command AContext Window=32k2026.06 | 79.6 | — | |
| GPT-4o (11/20)Context Window=32k2026.06 | 78.6 | — | |
| GRPO (Ts : 1.2)Sampling temperature (Ts)=1.2, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 77.4 | 90.6 | |
| TAMPOMaximum response length=6k tokens, Training algorithm=TAMPO2026.02 | 76.8 | 91 | |
| GRPO (Ts : 0.9 -> 1.5)Sampling temperature (Ts)=0.9 -> 1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 76.6 | 89.8 | |
| DS-Qwen-1.5BMaximum response length=6k tokens2026.02 | 76.2 | 89.2 | |
| GRPO (Ts : 1.5)Sampling temperature (Ts)=1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 75.4 | 90.8 | |
| GRPO (Ts : 0.9)Sampling temperature (Ts)=0.9, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 75.2 | 91 |