Mathematical Reasoning on OlympiadBench (Acc, Tokens, SuperBPE, Deltas)
66.8AccuracyShorthand for Thought (SFT)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Shorthand for Thought (SFT)Base Model=QWEN-3-30B-A3B, Training Condition=SFT2026.04 | 66.8 | 12,372 | 11.7 | 2.3 | -7 | |
| QWEN-3-30B-A3B BaselineBase Model=QWEN-3-30B-A3B, Training Condition=Baseline2026.04 | 65.3 | 13,309 | — | — | — | |
| Shorthand for Thought (SFT)Base Model=QwQ-32B, Training Condition=SFT2026.04 | 56.5 | 9,256 | 11.7 | 1.3 | -3 | |
| QwQ-32B BaselineBase Model=QwQ-32B, Training Condition=Baseline2026.04 | 55.8 | 9,584 | — | — | — | |
| DeepSeek-R1-Llama-70B-Distill BaselineBase Model=DeepSeek-R1-Llama-70B-Distill, Training Condition=Baseline2026.04 | 51 | 6,373 | — | — | — | |
| Shorthand for Thought (SFT)Base Model=DeepSeek-R1-Llama-70B-Distill, Training Condition=SFT2026.04 | 49.5 | 6,370 | 11.4 | -2.9 | -0.1 |