Mathematical Reasoning on MATH (Accuracy, Throughput, Speedup)
10.86SpeedupFast-Dual
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Fast-DualModel=Dream 7B, nu=12026.03 | 10.86 | 36.06 | 191.56 | — | — | — | |
| DyLLMModel=Dream 7B, Threshold (τ)=0.995, nu=12026.03 | 8.07 | 43.8 | 142.34 | — | — | — | |
| DyLLMModel=Dream 7B, Threshold (τ)=0.9975, nu=12026.03 | 7.4 | 45.12 | 130.57 | — | — | — | |
| DyLLMModel=LLaDA 8B, Threshold (τ)=0.99, nu=12026.03 | 6.71 | 38.08 | 106.06 | — | — | — | |
| Draft ModelModel=Qwen3-32B (Draft-1.7B)2026.06 | 6.42 | 65 | — | — | — | — | |
| DyLLMModel=LLaDA 8B, Threshold (τ)=0.995, nu=12026.03 | 6.13 | 38.68 | 96.98 | — | — | — | |
| ES-dLLMBase Model=Dream-7B-Base, Few-shot=42026.03 | 6 | — | 167.03 | 39.6 | — | — | |
| Fast-DualModel=LLaDA 8B, nu=12026.03 | 5.9 | 32.36 | 93.26 | — | — | — | |
| Draft ModelModel=DeepSeek R1-70B (Draft-8B)2026.06 | 4.53 | 53 | — | — | — | — | |
| STDecBackbone=LLaDA-8B-Instruct2026.04 | 4.46 | — | 53.05 | 39.86 | — | — | |
| STDecBase Model=Dream-7B-Instruct, Decoding Strategy=STDec2026.04 | 4.35 | — | 55.15 | 44.64 | — | — | |
| LLaDA + LocalLeapBackbone=LLaDA-8B-Instruct2026.04 | 4.2 | — | 49.93 | 39.66 | — | — | |
| LocalLeapBase Model=Dream-7B-Instruct, Decoding Strategy=LocalLeap2026.04 | 4.13 | — | 52.35 | 44.26 | — | — | |
| Fast-PrefixModel=LLaDA 8B, nu=12026.03 | 3.93 | 33.22 | 62.12 | — | — | — | |
| Fast-PrefixModel=Dream 7B, nu=12026.03 | 3.35 | 36.48 | 59.16 | — | — | — | |
| LLaDA + Fast-dLLM (Parallel)Backbone=LLaDA-8B-Instruct2026.04 | 3.28 | — | 39 | 41 | — | — | |
| Fast-dLLM (Parallel)Base Model=Dream-7B-Instruct, Decoding Strategy=Fast-dLLM (Parallel)2026.04 | 3.18 | — | 40.26 | 44.16 | — | — | |
| DualCacheBase Model=Dream-7B-Base, Few-shot=42026.03 | 3.1 | — | 86.58 | 40.1 | — | — | |
| MTP-D ensembleLoop strategy=4 to 8, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.972 | — | — | — | — | — | |
| PARSE + EAGLE3Model=Qwen3-235B-A22B, Draft Model=Qwen3-8B, Decoding Strategy=PARSE with EAGLE3 on both draft and target, Precision=FP8, Inference Framework=SGLang2026.05 | 2.83 | — | 272.2 | — | — | — | |
| MTP-DLoop strategy=4 to 8, Training protocol=Continued pre-training, Training data size=350B tokens2026.03 | 2.818 | — | — | — | — | — | |
| MTP-DLoop strategy=4 to 8, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.8 | — | — | — | — | — | |
| dLLM-CacheModel=Dream 7B, nu=12026.03 | 2.76 | 44.98 | 48.76 | — | — | — | |
| MTP-DLoop strategy=1 to 8, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.741 | — | — | — | — | — | |
| MTP-DLoop strategy=4 to 16, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.555 | — | — | — | — | — | |
| MTPLoop strategy=4 to 8, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.547 | — | — | — | — | — | |
| MTP-DLoop strategy=4 to 8 to 16, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.465 | — | — | — | — | — | |
| MTP-DLoop strategy=4 to 8, Training protocol=Training Free, Training data size=70B tokens2026.03 | 2.358 | — | — | — | — | — | |
| MTPLoop strategy=1 to 8, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.343 | — | — | — | — | — | |
| MTP-DLoop strategy=4, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.33 | — | — | — | — | — | |
| dLLM-CacheModel=LLaDA 8B, nu=12026.03 | 2.31 | 24.7 | 36.56 | — | — | — | |
| PARSEModel=Qwen3-235B-A22B, Draft Model=Qwen3-8B, Decoding Strategy=Full and partial verify, Precision=FP8, Inference Framework=SGLang2026.05 | 2.24 | 75.6 | 215.7 | — | — | — | |
| Specreason+RPOModel=Qwen3-32B (Draft-1.7B)2026.06 | 2.21 | 76 | — | — | — | — | |
| Specreason+SRModel=Qwen3-32B (Draft-1.7B)2026.06 | 2.18 | 76 | — | — | — | — | |
| MTP-DLoop strategy=1 to 16, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 2.155 | — | — | — | — | — | |
| SpecreasonModel=Qwen3-32B (Draft-1.7B)2026.06 | 2.14 | 76 | — | — | — | — | |
| MTP-DLoop strategy=4 to 16, Training protocol=Training Free, Training data size=70B tokens2026.03 | 2.095 | — | — | — | — | — | |
| SafeSpecModel=Qwen3-32B (Draft-1.7B)2026.06 | 2.06 | 78 | — | — | — | — | |
| Half-StepBase Model=Dream-7B-Instruct, Decoding Strategy=Half-Step2026.04 | 2 | — | 25.31 | 39.4 | — | — | |
| LLaDA + Half-StepBackbone=LLaDA-8B-Instruct2026.04 | 1.99 | — | 23.64 | 39.44 | — | — | |
| Specreason+RPOModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1.98 | 61 | — | — | — | — | |
| Specreason+SRModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1.94 | 60 | — | — | — | — | |
| SpecreasonModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1.81 | 62 | — | — | — | — | |
| PARSETarget Model=GLM-4.7-FP8, Draft Model=Qwen3.5-9B2026.05 | 1.79 | 78.4 | 126 | — | — | — | |
| SafeSpecModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1.76 | 61 | — | — | — | — | |
| MTP-DLoop strategy=1 to 8, Training protocol=Training Free, Training data size=70B tokens2026.03 | 1.726 | — | — | — | — | — | |
| Eagle3Model=Qwen3-235B-A22B, Decoding Strategy=Speculative decoding, Precision=FP8, Inference Framework=SGLang2026.05 | 1.71 | — | 164.3 | — | — | — | |
| Spec-DecodingModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1.7 | 65 | — | — | — | — | |
| MTP-DLoop strategy=1 to 16, Training protocol=Training Free, Training data size=70B tokens2026.03 | 1.443 | — | — | — | — | — | |
| Spec-DecodingModel=Qwen3-32B (Draft-1.7B)2026.06 | 1.38 | 78 | — | — | — | — | |
| LLaDA + dKV-CacheBackbone=LLaDA-8B-Instruct2026.04 | 1.24 | — | 14.73 | 40.9 | — | — | |
| dKV-CacheBase Model=Dream-7B-Instruct, Decoding Strategy=dKV-Cache2026.04 | 1.23 | — | 15.56 | 44.04 | — | — | |
| MTP-DLoop strategy=1, Training protocol=Continued pre-training, Training data size=70B tokens2026.03 | 1.188 | — | — | — | — | — | |
| RPOModel=Qwen3-32B (Draft-1.7B)2026.06 | 1.16 | 80 | — | — | — | — | |
| RPOModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1.12 | 65 | — | — | — | — | |
| SRModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1.09 | 64 | — | — | — | — | |
| SRModel=Qwen3-32B (Draft-1.7B)2026.06 | 1.07 | 75 | — | — | — | — | |
| OriginalModel=LLaDA 8B, nu=12026.03 | 1 | 33.22 | 15.81 | — | — | — | |
| OriginalModel=Dream 7B, nu=12026.03 | 1 | 37.6 | 17.64 | — | — | — | |
| DreamBase Model=Dream-7B-Base, Few-shot=42026.03 | 1 | — | 27.98 | 40.9 | — | — | |
| Vanilla DreamBase Model=Dream-7B-Instruct, Decoding Strategy=Vanilla2026.04 | 1 | — | 12.67 | 44.64 | — | — | |
| Vanilla LLaDABackbone=LLaDA-8B-Instruct2026.04 | 1 | — | 11.9 | 40.38 | — | — | |
| No DefenseModel=Qwen3-32B (Draft-1.7B)2026.06 | 1 | 79 | — | — | — | — | |
| No DefenseModel=DeepSeek R1-70B (Draft-8B)2026.06 | 1 | 64 | — | — | — | — | |
| SafeDecodingModel=Qwen3-32B (Draft-1.7B)2026.06 | 0.82 | 74 | — | — | — | — | |
| SafeDecodingModel=DeepSeek R1-70B (Draft-8B)2026.06 | 0.78 | 63 | — | — | — | — | |
| SecDecodingModel=Qwen3-32B (Draft-1.7B)2026.06 | 0.61 | 75 | — | — | — | — | |
| SecDecodingModel=DeepSeek R1-70B (Draft-8B)2026.06 | 0.57 | 60 | — | — | — | — | |
| AdaMoESparsity Level=Mid Sparsity, Null=1282026.05 | — | 41.68 | — | — | — | — | |
| AdaMoESparsity Level=High Sparsity, Null=2562026.05 | — | 34.06 | — | — | — | — | |
| AdamWModel=MOE-68B-A3B, Training Tokens=700B, Evaluation Mode=Chain of Thought (CoT)2026.05 | — | — | — | — | 0.1648 | — | |
| BEAMSparsity Level=Mid Sparsity, beta=0.012026.05 | — | 55.16 | — | — | — | — | |
| BEAMSparsity Level=High Sparsity, beta=0.12026.05 | — | 55.44 | — | — | — | — | |
| BEAMSparsity Level=Extreme Sparsity, beta=1.02026.05 | — | 52.2 | — | — | — | — | |
| D2FBackbone=LLaDA-Instruct2026.04 | — | 28.7 | 2.38 | — | — | — | |
| d3LLMInstruction Tuning=Dream-Instruct2026.04 | — | 38.2 | 3.92 | — | — | — | |
| d3LLMBackbone=LLaDA-Instruct2026.04 | — | 30.4 | 5.74 | — | — | — | |
| dParallelInstruction Tuning=Dream-Instruct2026.04 | — | 38.7 | 2.94 | — | — | — | |
| dParallelBackbone=LLaDA-Instruct2026.04 | — | 30.2 | 3.17 | — | — | — | |
| DreamInstruction Tuning=Dream-Instruct2026.04 | — | 39.6 | 1 | — | — | — | |
| Fast-dLLMInstruction Tuning=Dream-Instruct2026.04 | — | 38.3 | 1.78 | — | — | — | |
| Fast-dLLMBackbone=LLaDA-Instruct2026.04 | — | 30.8 | 1.97 | — | — | — | |
| Fast-dLLM-v2Instruction Tuning=Dream-Instruct2026.04 | — | 48.7 | 2.61 | — | — | — | |
| GLM-4.7-FP8Mode=Standalone2026.05 | — | 79.6 | 70.2 | — | — | — | |
| LLaDABackbone=LLaDA-Instruct2026.04 | — | 32.2 | 1 | — | — | — | |
| Mellum 2Parameters=2.5B/12B2026.05 | — | — | — | — | — | 10 | |
| MoE-DynamicSparsity Level=Mid Sparsity, phi=0.32026.05 | — | 54.28 | — | — | — | — | |
| MoE-DynamicSparsity Level=High Sparsity, phi=0.12026.05 | — | 40.1 | — | — | — | — | |
| MONAModel=MOE-68B-A3B, Training Tokens=700B, Evaluation Mode=Chain of Thought (CoT)2026.05 | — | — | — | — | 0.1594 | — | |
| MuonModel=MOE-68B-A3B, Training Tokens=700B, Evaluation Mode=Chain of Thought (CoT)2026.05 | — | — | — | — | 0.1684 | — | |
| OLMo-3-7BParameters=7B2026.05 | — | — | — | — | — | 18.7 | |
| Qwen2.5-7BParameters=7B2026.05 | — | — | — | — | — | 24.6 | |
| Qwen3-235B-A22BModel=Qwen3-235B-A22B, Decoding Strategy=Plain autoregressive, Precision=FP8, Inference Framework=SGLang2026.05 | — | 70.8 | 96.1 | — | — | — | |
| Qwen3-30B-A3BSparsity Level=Base, K=82026.05 | — | 58.28 | — | — | — | — | |
| Qwen3-4BParameters=4B2026.05 | — | — | — | — | — | 27.7 | |
| Qwen3-8BModel=Qwen3-8B, Decoding Strategy=Standalone2026.05 | — | 76.4 | — | — | — | — | |
| Qwen3.5-4BParameters=4B2026.05 | — | — | — | — | — | 25.3 | |
| Qwen3.5-9BMode=Standalone2026.05 | — | 72.8 | — | — | — | — | |
| Top-K PruningSparsity Level=Mid Sparsity, K=42026.05 | — | 48.76 | — | — | — | — | |
| Top-K PruningSparsity Level=High Sparsity, K=22026.05 | — | 0.68 | — | — | — | — |