Mathematical Reasoning on GSM8K standard (test) (Acc., Tokens, CR., TE.)
94.16AccuracyTokenSkip
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TokenSkipModel=Qwen2.5-14B, Compression Strength=Low--2026.01 | 94.16 | 269.52 | 86 | 34.94 | |
| CtrlCoTModel=Qwen2.5-14B, Compression Strength=Low--2026.01 | 94.01 | 261.18 | 83 | 35.99 | |
| CtrlCoTModel=Qwen2.5-14B, Compression Strength=Low-2026.01 | 93.63 | 243.9 | 78 | 38.39 | |
| TokenSkipModel=Qwen2.5-14B, Compression Strength=Low-2026.01 | 93.48 | 250.83 | 80 | 37.27 | |
| OriginalModel=Qwen2.5-14B, Compression Strength=None2026.01 | 93.03 | 313.94 | 100 | 29.63 | |
| OriginalModel=Qwen2.5-7B, Compression Strength=None2026.01 | 91.58 | 299.22 | 100 | 3.27 | |
| CtrlCoTModel=Qwen2.5-7B, Compression Strength=Low--2026.01 | 91.13 | 252.92 | 85 | 36.03 | |
| TokenSkipModel=Qwen2.5-7B, Compression Strength=Low--2026.01 | 90.6 | 262.53 | 88 | 34.51 | |
| TruncationModel=Qwen2.5-14B, Compression Strength=Low--2026.01 | 90.3 | 311.37 | 99 | 29 | |
| TokenSkipModel=Qwen2.5-7B, Compression Strength=Low-2026.01 | 90.14 | 243.77 | 81 | 36.98 | |
| TruncationModel=Qwen2.5-7B, Compression Strength=Low--2026.01 | 90.07 | 296.73 | 99 | 30.35 | |
| CtrlCoTModel=Qwen2.5-7B, Compression Strength=Low-2026.01 | 89.31 | 218.22 | 73 | 40.93 | |
| CF-DAPOModel=Qwen2.5-3B, Weighting Mode=Counterfactual, Training Steps=5002026.02 | 86.7 | — | — | — | |
| TruncationModel=Qwen2.5-14B, Compression Strength=Low-2026.01 | 86.2 | 305.95 | 97 | 28.17 | |
| TruncationModel=Qwen2.5-7B, Compression Strength=Low-2026.01 | 85.97 | 292.77 | 98 | 29.37 | |
| Random weightingModel=Qwen2.5-3B, Weighting Mode=Random, Training Steps=5002026.02 | 85.8 | — | — | — | |
| DAPOModel=Qwen2.5-3B, Weighting Mode=Vanilla, Training Steps=5002026.02 | 85.6 | — | — | — | |
| Inverted weightingModel=Qwen2.5-3B, Weighting Mode=Inverted, Training Steps=5002026.02 | 85.3 | — | — | — | |
| CF-DAPOModel=Qwen3-1.7B, Weighting Mode=Counterfactual, Training Steps=5002026.02 | 84.3 | — | — | — | |
| DAPOModel=Qwen3-1.7B, Weighting Mode=Vanilla, Training Steps=5002026.02 | 83.4 | — | — | — | |
| Random weightingModel=Qwen3-1.7B, Weighting Mode=Random, Training Steps=5002026.02 | 82.5 | — | — | — | |
| Inverted weightingModel=Qwen3-1.7B, Weighting Mode=Inverted, Training Steps=5002026.02 | 81.9 | — | — | — | |
| CF-DAPOModel=Llama3.2-3B, Weighting Mode=Counterfactual, Training Steps=5002026.02 | 78.9 | — | — | — | |
| DAPOModel=Llama3.2-3B, Weighting Mode=Vanilla, Training Steps=5002026.02 | 78.2 | — | — | — | |
| Random weightingModel=Llama3.2-3B, Weighting Mode=Random, Training Steps=5002026.02 | 78.1 | — | — | — | |
| Inverted weightingModel=Llama3.2-3B, Weighting Mode=Inverted, Training Steps=5002026.02 | 76.4 | — | — | — | |
| Phi-4-mini-InstPrecision=BF162026.06 | 75.66 | — | — | — | |
| Llama-3.1-8B-InstPrecision=BF162026.06 | 73.54 | — | — | — | |
| Qwen3-8BPrecision=BF162026.06 | 56.1 | — | — | — | |
| Qwen3.5-35B-A3BPrecision=BF162026.06 | 54.44 | — | — | — | |
| Qwen2.5-3B-InstPrecision=BF162026.06 | 45.56 | — | — | — | |
| Mistral-7B-InstPrecision=BF162026.06 | 43.37 | — | — | — | |
| Llama-3.1-8B (base)Precision=BF162026.06 | 25.63 | — | — | — |