Inference Efficiency on GSM8K, MATH500, MBPP+, HumanEval+ Average
9.34Avg. TPFMBD-LLaDA2-Mini-DMax
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| MBD-LLaDA2-Mini-DMaxHardware=two H100 GPUs, Tensor Parallelism=TP=2, Decoding=single-sample2026.06 | 9.34 | 169.16 | 11.2 | 1.58 | 926.67 | 79.19 | |
| LLaDA2-Mini-DMaxHardware=two H100 GPUs, Tensor Parallelism=TP=2, Decoding=single-sample2026.06 | 6.35 | 83 | 9.02 | 1.28 | 779.49 | 50.73 | |
| MBD-LLaDA2-MiniHardware=two H100 GPUs, Tensor Parallelism=TP=2, Decoding=single-sample2026.06 | 6.19 | 78.39 | 8.78 | 1.24 | 745.92 | 44.24 | |
| LLaDA2-MiniHardware=two H100 GPUs, Tensor Parallelism=TP=2, Decoding=single-sample2026.06 | 3.47 | — | 7.07 | 1 | 517.16 | — |