Reasoning on GSM8K (maj@2, maj@4, maj@8)
41.17Accuracy (maj@4)DLE (top-p&top-k)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DLE (top-p&top-k)Algorithm=DLE, Sampling Strategy=top-p&top-k, Model=Qwen2.5-0.5B-Instruct2026.04 | 41.17 | 34.8 | 44.43 | |
| DLE (ε-sampling)Algorithm=DLE, Sampling Strategy=ε-sampling, Model=Qwen2.5-0.5B-Instruct2026.04 | 40.64 | 34.57 | 44.05 | |
| DLE (min-p)Algorithm=DLE, Sampling Strategy=min-p, Model=Qwen2.5-0.5B-Instruct2026.04 | 39.58 | 33.74 | 43.97 | |
| Beam searchAlgorithm=Beam search, Model=Qwen2.5-0.5B-Instruct2026.04 | 36.47 | 35.1 | 36.92 | |
| Diverse beam searchAlgorithm=Diverse beam search, Model=Qwen2.5-0.5B-Instruct2026.04 | 36.47 | — | 41.39 | |
| Self-consistencyAlgorithm=Self-consistency, Temperature (tau)=0.6, Model=Qwen2.5-0.5B-Instruct2026.04 | 36.24 | 29.57 | 41.32 | |
| Self-consistency (min-p)Algorithm=Self-consistency, Sampling Strategy=min-p, Model=Qwen2.5-0.5B-Instruct2026.04 | 33.51 | 26.46 | 39.04 | |
| Self-consistency (ε-sampling)Algorithm=Self-consistency, Sampling Strategy=ε-sampling, Model=Qwen2.5-0.5B-Instruct2026.04 | 33.51 | 26.84 | 40.94 | |
| Self-consistency (top-p&top-k)Algorithm=Self-consistency, Sampling Strategy=top-p&top-k, Model=Qwen2.5-0.5B-Instruct2026.04 | 31.84 | 25.47 | 38.74 | |
| DeepConfAlgorithm=DeepConf, Model=Qwen2.5-0.5B-Instruct2026.04 | 26.91 | 21.68 | 33.36 | |
| Self-consistencyAlgorithm=Self-consistency, Temperature (tau)=1, Model=Qwen2.5-0.5B-Instruct2026.04 | 25.32 | 19.26 | 31.61 | |
| Self-certaintyAlgorithm=Self-certainty, Model=Qwen2.5-0.5B-Instruct2026.04 | 25.32 | 20.09 | 32.22 |