Mathematical Reasoning on GSM8K (maj@2, maj@4, maj@8)
64.52Accuracy (maj@2)DLE (top-p+top-k)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DLE (top-p+top-k)Model=Llama3.2-3B-Instruct, Sampling method=top-p+top-k, Decoding algorithm=Distinct Leaf Enumeration2026.04 | 64.52 | 69.45 | 71.49 | |
| DLE (ε-sampling)-RANDBRANCHModel=Llama3.2-3B-Instruct, Sampling method=ε-sampling, Branching strategy=RANDBRANCH2026.04 | 64.52 | 68.31 | 72.33 | |
| DLE (min-p)Model=Llama3.2-3B-Instruct, Sampling method=min-p, Decoding algorithm=Distinct Leaf Enumeration2026.04 | 64.44 | 68.84 | 71.57 | |
| DLE (ε-sampling)-PROBFIRSTModel=Llama3.2-3B-Instruct, Sampling method=ε-sampling, Branching strategy=PROBFIRST2026.04 | 64.44 | 68.92 | 71.65 | |
| DLE (ε-sampling)-DIVFIRSTModel=Llama3.2-3B-Instruct, Sampling method=ε-sampling, Branching strategy=DIVFIRST2026.04 | 62.62 | 67.4 | 69.52 | |
| Self-consistency (min-p)Model=Llama3.2-3B-Instruct, Sampling method=min-p2026.04 | 59.89 | 68.39 | 73.69 | |
| Self-consistency (ε-sampling)Model=Llama3.2-3B-Instruct, Sampling method=ε-sampling2026.04 | 58.38 | 62.33 | 69.14 | |
| Self-consistency (top-p+top-k)Model=Llama3.2-3B-Instruct, Sampling method=top-p+top-k2026.04 | 56.25 | 65.81 | 72.48 | |
| Self-consistencyModel=Llama3.2-3B-Instruct2026.04 | 43.29 | 54.36 | 65.58 |