Mathematical Reasoning on GSM8K (128 samples subset)
96.2Top-1 AccuracyBeam-Search
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Beam-SearchSelection=Argmax, Scoring=PRM, Base Model=Qwen2.5-7B-Instruct2025.10 | 96.2 | — | — | |
| PFSelection=Argmax, Scoring=PRM, Base Model=Qwen2.5-7B-Instruct2025.10 | 96.2 | — | — | |
| Best-of-NSelection=Argmax, Scoring=ORM, Base Model=Qwen2.5-7B-Instruct2025.10 | 96 | — | — | |
| ePFSelection=Argmax, Scoring=PRM, Base Model=Qwen2.5-7B-Instruct2025.10 | 95.8 | — | — | |
| Self-ConsistencySelection=MV, Base Model=Qwen2.5-7B-Instruct2025.10 | 94.8 | — | — | |
| PFSelection=Argmax, Scoring=PRM, Base Model=Qwen2.5-1.5B-Instruct2025.10 | 93.75 | — | — | |
| ePFSelection=Argmax, Scoring=PRM, Base Model=Qwen2.5-1.5B-Instruct2025.10 | 93.75 | — | — | |
| Base SamplingBase Model=Qwen2.5-7B-Instruct2025.10 | 93.4 | — | — | |
| Best-of-NSelection=Argmax, Scoring=ORM, Base Model=Qwen2.5-1.5B-Instruct2025.10 | 92.97 | — | — | |
| Beam-SearchSelection=Argmax, Scoring=PRM, Base Model=Qwen2.5-1.5B-Instruct2025.10 | 91.4 | — | — | |
| Self-ConsistencySelection=MV, Base Model=Qwen2.5-1.5B-Instruct2025.10 | 82.03 | — | — | |
| Base SamplingBase Model=Qwen2.5-1.5B-Instruct2025.10 | 67.19 | — | — | |
| LLaMA-2Model Type=Autoregressive Baseline, Parameters=7B, Avg NFE=N/A2025.12 | — | 10.69 | 11.6 | |
| LLaMA-3Model Type=Autoregressive Baseline, Parameters=8B, Variant=Instruct, Avg NFE=N/A2025.12 | — | 61.11 | 68.99 | |
| Method 1 (Token Injection)Initialization=Context-Aware, Mechanism=Hard token injection, Avg NFE=51.72025.12 | — | 22.66 | 26.56 | |
| Method 2 (Embedding Interp.)Initialization=Context-Aware, Mechanism=Soft embedding interpolation, Avg NFE=56.42025.12 | — | 25 | 26.56 | |
| Original Fast-dLLMModel Type=Diffusion Baseline, Avg NFE=79.122025.12 | — | 35.94 | 78.12 |