Mathematical Reasoning on MATH 500 (Accuracy, Pass@1)
95.2AccuracySegment Selective SFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-7B, SFT Training Strategy=Segment Selective SFT, Decoding Strategy=Greedy Decoding2026.01 | 95.2 | — | |
| Full CoT SFTModel Backbone=R1-Distill-Qwen-7B, SFT Training Strategy=Full CoT SFT, Decoding Strategy=Greedy Decoding2026.01 | 91.2 | — | |
| R1-Distill-Qwen-7BModel Backbone=R1-Distill-Qwen-7B, SFT Training Strategy=None/Baseline, Decoding Strategy=Greedy Decoding2026.01 | 85.8 | — | |
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-1.5B, SFT Training Strategy=Segment Selective SFT, Decoding Strategy=Greedy Decoding2026.01 | 82.4 | — | |
| Full CoT SFTModel Backbone=R1-Distill-Qwen-1.5B, SFT Training Strategy=Full CoT SFT, Decoding Strategy=Greedy Decoding2026.01 | 80.8 | — | |
| Full CoT SFTModel Backbone=Qwen2.5-7B-Instruct, SFT Training Strategy=Full CoT SFT, Decoding Strategy=Greedy Decoding2026.01 | 77.2 | — | |
| Qwen2.5-7B-InstructModel Backbone=Qwen2.5-7B-Instruct, SFT Training Strategy=None/Baseline, Decoding Strategy=Greedy Decoding2026.01 | 77 | — | |
| Segment Selective SFTModel Backbone=Qwen2.5-7B-Instruct, SFT Training Strategy=Segment Selective SFT, Decoding Strategy=Greedy Decoding2026.01 | 76.6 | — | |
| R1-Distill-Qwen-1.5BModel Backbone=R1-Distill-Qwen-1.5B, SFT Training Strategy=None/Baseline, Decoding Strategy=Greedy Decoding2026.01 | 70.6 | — | |
| Full CoT SFTModel Backbone=R1-Distill-Qwen-1.5B, SFT Training Strategy=Full CoT SFT, Decoding Strategy=Temperature Sampling2026.01 | — | 85.1 | |
| Full CoT SFTModel Backbone=R1-Distill-Qwen-7B, SFT Training Strategy=Full CoT SFT, Decoding Strategy=Temperature Sampling2026.01 | — | 93.6 | |
| Full CoT SFTModel Backbone=Qwen2.5-7B-Instruct, SFT Training Strategy=Full CoT SFT, Decoding Strategy=Temperature Sampling2026.01 | — | 77 | |
| Qwen2.5-7B-InstructModel Backbone=Qwen2.5-7B-Instruct, SFT Training Strategy=None/Baseline, Decoding Strategy=Temperature Sampling2026.01 | — | 75.7 | |
| R1-Distill-Qwen-1.5BModel Backbone=R1-Distill-Qwen-1.5B, SFT Training Strategy=None/Baseline, Decoding Strategy=Temperature Sampling2026.01 | — | 84 | |
| R1-Distill-Qwen-7BModel Backbone=R1-Distill-Qwen-7B, SFT Training Strategy=None/Baseline, Decoding Strategy=Temperature Sampling2026.01 | — | 92.8 | |
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-1.5B, SFT Training Strategy=Segment Selective SFT, Decoding Strategy=Temperature Sampling2026.01 | — | 85.1 | |
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-7B, SFT Training Strategy=Segment Selective SFT, Decoding Strategy=Temperature Sampling2026.01 | — | 93.3 | |
| Segment Selective SFTModel Backbone=Qwen2.5-7B-Instruct, SFT Training Strategy=Segment Selective SFT, Decoding Strategy=Temperature Sampling2026.01 | — | 77.1 |