Mathematics Reasoning on GSM8K (Accuracy)
91.3AccuracySDAR-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| SDAR-8BScale=8B, Tokens=55B, FLOPs=2640, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 91.3 | |
| SDAR-4BScale=4B, Tokens=55B, FLOPs=1320, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 89.9 | |
| Fisher Merging w/ HARCMerging Strategy=Fisher Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 87.99 | |
| IndividualModel Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 87.97 | |
| Fisher MergingMerging Strategy=Fisher Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 87.95 | |
| Weight Averaging w/ HARCMerging Strategy=Weight Averaging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 87.91 | |
| Weight AveragingMerging Strategy=Weight Averaging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 87.86 | |
| TIES-Merging w/ HARCMerging Strategy=TIES-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 87.76 | |
| DAREMerging Strategy=DARE, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 87.65 | |
| TIES-MergingMerging Strategy=TIES-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 87.64 | |
| DARE w/ HARCMerging Strategy=DARE, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 87.61 | |
| OPDLM-4BScale=4B, Tokens=0.076B, FLOPs=2.4, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 87.6 | |
| WUDI-Merging w/ HARCMerging Strategy=WUDI-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 87.53 | |
| WUDI-MergingMerging Strategy=WUDI-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 87.34 | |
| OPDLM-8BScale=8B, Tokens=0.066B, FLOPs=4.2, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 87.1 | |
| Fast-dLLM-v2-7BScale=8B, Tokens=1B, FLOPs=42, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 83.7 | |
| Dream-7BScale=8B, Tokens=580B, FLOPs=24360, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 81 | |
| LLaDA-8BScale=8B, Tokens=1500B, FLOPs=72000, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 78.6 |