Mathematical Reasoning on AMC 23 (Accuracy, Pass@1)
90AccuracyFull CoT SFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Full CoT SFTModel Backbone=R1-Distill-Qwen-7B, Decoding Strategy=Greedy Decoding2026.01 | 90 | — | |
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-7B, Decoding Strategy=Greedy Decoding2026.01 | 90 | — | |
| R1-Distill-Qwen-7BModel Backbone=R1-Distill-Qwen-7B, Decoding Strategy=Greedy Decoding2026.01 | 85 | — | |
| Full CoT SFTModel Backbone=R1-Distill-Qwen-1.5B, Decoding Strategy=Greedy Decoding2026.01 | 70 | — | |
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-1.5B, Decoding Strategy=Greedy Decoding2026.01 | 60 | — | |
| Segment Selective SFTModel Backbone=Qwen2.5-7B-Instruct, Decoding Strategy=Greedy Decoding2026.01 | 57.5 | — | |
| Full CoT SFTModel Backbone=Qwen2.5-7B-Instruct, Decoding Strategy=Greedy Decoding2026.01 | 55 | — | |
| R1-Distill-Qwen-1.5BModel Backbone=R1-Distill-Qwen-1.5B, Decoding Strategy=Greedy Decoding2026.01 | 52.5 | — | |
| Qwen2.5-7B-InstructModel Backbone=Qwen2.5-7B-Instruct, Decoding Strategy=Greedy Decoding2026.01 | 50 | — | |
| Full CoT SFTModel Backbone=R1-Distill-Qwen-1.5B, Decoding Strategy=Temperature Sampling2026.01 | — | 74.7 | |
| R1-Distill-Qwen-1.5BModel Backbone=R1-Distill-Qwen-1.5B, Decoding Strategy=Temperature Sampling2026.01 | — | 73.3 | |
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-1.5B, Decoding Strategy=Temperature Sampling2026.01 | — | 75.9 |