Mathematical Reasoning on GSM8K (Acc, TPF, TPS, AUP Score)
2.04Time Per First Token (TPF)LLaDA-2.0-mini
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaDA-2.0-miniZero-shot=true, Batch size=1, Decoding threshold=0.95, Generation length=2048, Hardware=2x H200 GPUs, Framework=dInFer [54], Prompting strategy=chain-of-thought [80]2026.04 | 2.04 | 512 | 92.6 | 340 | |
| Uniform Diffusion TrainingZero-shot=true, Batch size=1, Generation length=2048, Hardware=2x H200 GPUs, Framework=dInFer [54], Prompting strategy=chain-of-thought [80]2026.04 | 2.26 | 493 | 68.7 | 0 | |
| Hierarchical DecodingZero-shot=true, Batch size=1, Decoding threshold=0.2, Generation length=2048, Hardware=2x H200 GPUs, Framework=dInFer [54], Prompting strategy=chain-of-thought [80]2026.04 | 2.44 | 577 | 91.6 | 357 | |
| dParallel SFTZero-shot=true, Batch size=1, Generation length=2048, Hardware=2x H200 GPUs, Framework=dInFer [54], Prompting strategy=chain-of-thought [80]2026.04 | 2.79 | 721 | 92.3 | 395 | |
| DMax-MathZero-shot=true, Batch size=1, Decoding threshold=0.5, Generation length=2048, Hardware=2x H200 GPUs, Framework=dInFer [54], Prompting strategy=chain-of-thought [80]2026.04 | 5.48 | 1,258 | 92.1 | 557 |