Multi-task Language Understanding on MMLU-Pro (maj@2, maj@4, maj@8)
54.88Accuracy (maj@2)DLE (ε-sampling)-RANDBRANCH
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DLE (ε-sampling)-RANDBRANCHBackbone=Qwen2.5-14B-Instruct, Sampling Strategy=ε-sampling, Branching Algorithm=RANDBRANCH2026.04 | 54.88 | 56.52 | 58.29 | |
| DLE (ε-sampling)-DIVFIRSTBackbone=Qwen2.5-14B-Instruct, Sampling Strategy=ε-sampling, Branching Algorithm=DIVFIRST2026.04 | 54.66 | 56.4 | 58.18 | |
| DLE (top-p+top-k)Backbone=Qwen2.5-14B-Instruct, Sampling Strategy=top-p+top-k2026.04 | 54.5 | 56.27 | 58.06 | |
| DLE (min-p)Backbone=Qwen2.5-14B-Instruct, Sampling Strategy=min-p2026.04 | 54.34 | 56.13 | 58.23 | |
| DLE (ε-sampling)-PROBFIRSTBackbone=Qwen2.5-14B-Instruct, Sampling Strategy=ε-sampling, Branching Algorithm=PROBFIRST2026.04 | 54.33 | 56.38 | 58.3 | |
| Self-consistency (ε-sampling)Backbone=Qwen2.5-14B-Instruct, Sampling Strategy=ε-sampling2026.04 | 52.94 | 56.48 | 58.56 | |
| Self-consistency (top-p+top-k)Backbone=Qwen2.5-14B-Instruct, Sampling Strategy=top-p+top-k2026.04 | 52.68 | 55.31 | 57.74 | |
| Self-consistency (min-p)Backbone=Qwen2.5-14B-Instruct, Sampling Strategy=min-p2026.04 | 52.68 | 55.31 | 57.74 | |
| Self-consistencyBackbone=Qwen2.5-14B-Instruct, Sampling Strategy=Standard2026.04 | 52.01 | 55.5 | 58.39 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=min-p2026.04 | 20.71 | 22.17 | 22.85 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=top-p+top-k2026.04 | 20.62 | 22.17 | 22.86 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling, Branching Strategy=DIVFIRST2026.04 | 19.76 | 21.23 | 22.15 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling, Branching Strategy=RANDBRANCH2026.04 | 19.75 | 20.9 | 21.97 | |
| DLEModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling, Branching Strategy=PROBFIRST2026.04 | 19.61 | 21.13 | 22.95 | |
| Self-consistencyModel=Llama3.2-1B-Instruct, Sampling Strategy=ε-sampling2026.04 | 18.87 | 20.69 | 21.82 | |
| Self-consistencyModel=Llama3.2-1B-Instruct, Sampling Strategy=min-p2026.04 | 17.73 | 20.09 | 21.29 | |
| Self-consistencyModel=Llama3.2-1B-Instruct, Sampling Strategy=top-p+top-k2026.04 | 17.3 | 19.3 | 20.78 | |
| Self-consistencyModel=Llama3.2-1B-Instruct2026.04 | 13.73 | 16.71 | 19.02 |