Mathematical Reasoning on LMB Hard
46.2AccuracyQwen2.5-32B-Instruct + Bootcamp-SFT-RL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-32B-Instruct + Bootcamp-SFT-RLModel=Qwen2.5-32B-Instruct, Training Stage=Bootcamp-SFT-RL2025.08 | 46.2 | |
| DS-R1-Distilled-Qwen-32B + Bootcamp-RLModel=DS-R1-Distilled-Qwen-32B, Training Stage=Bootcamp-RL2025.08 | 43.7 | |
| DS-R1-Distilled-Qwen-32BModel=DS-R1-Distilled-Qwen-32B2025.08 | 36.8 | |
| Qwen2.5-32B-Instruct + Bootcamp-SFTModel=Qwen2.5-32B-Instruct, Training Stage=Bootcamp-SFT2025.08 | 33.2 | |
| Qwen2.5-32B-InstructModel=Qwen2.5-32B-Instruct2025.08 | 22 | |
| Qwen2.5-32B-Instruct + Bootcamp-RLModel=Qwen2.5-32B-Instruct, Training Stage=Bootcamp-RL2025.08 | 21.8 | |
| OPDLM-8BScale=8B, Tokens=0.066B, FLOPs=4.2, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 20 | |
| OPDLM-4BScale=4B, Tokens=0.076B, FLOPs=2.4, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 11.1 | |
| SDAR-8BScale=8B, Tokens=55B, FLOPs=2640, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 8.9 | |
| Fast-dLLM-v2-7BScale=8B, Tokens=1B, FLOPs=42, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 8.9 | |
| SDAR-4BScale=4B, Tokens=55B, FLOPs=1320, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 6.9 |