Mathematical Reasoning on GSM8K (maj@k)
33.21Accuracy (maj@2)DLE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=min-p2026.04 | 33.21 | 36.01 | 38.44 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling, Branching Strategy=PROBFIRST2026.04 | 33.21 | 36.01 | 38.06 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling, Branching Strategy=RANDBRANCH2026.04 | 33.13 | 34.8 | 38.44 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=top-p+top-k2026.04 | 33.06 | 36.01 | 38.06 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling, Branching Strategy=DIVFIRST2026.04 | 31.39 | 34.65 | 36.85 | |
| Self-consistencyModel=Llama3.2-1B-Instruct, Sampling Strategy=min-p2026.04 | 30.1 | 35.25 | 40.25 | |
| Self-consistencyModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling2026.04 | 27.75 | 33.89 | 39.95 | |
| Self-consistencyModel=Llama3.2-1B-Instruct, Sampling Strategy=top-p+top-k2026.04 | 27.6 | 33.21 | 39.65 | |
| Self-consistencyModel=Llama3.2-1B-Instruct2026.04 | 18.5 | 24.64 | 32.6 |