Commonsense Reasoning on Winogrande (Accuracy, Speedup)
0.8624AccuracyExp-k=2 (7.5, 2.5)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Exp-k=2 (7.5, 2.5)Schedule Type=Exponential, k=2, Thresholds (τ_high, τ_low)=(7.5, 2.5)2025.12 | 0.8624 | 1.23 | |
| Exp-k=2 (7.5, 0)Schedule Type=Exponential, k=2, Thresholds (τ_high, τ_low)=(7.5, 0)2025.12 | 0.8616 | 1.54 | |
| Cosine (7.5, 0)Schedule Type=Cosine, Thresholds (τ_high, τ_low)=(7.5, 0)2025.12 | 0.8596 | 1.32 | |
| Linear (7.5, 0)Schedule Type=Linear, Thresholds (τ_high, τ_low)=(7.5, 0)2025.12 | 0.8596 | 1.32 | |
| Cosine (7.5, 2.5)Schedule Type=Cosine, Thresholds (τ_high, τ_low)=(7.5, 2.5)2025.12 | 0.8591 | 1.23 | |
| Linear (7.5, 2.5)Schedule Type=Linear, Thresholds (τ_high, τ_low)=(7.5, 2.5)2025.12 | 0.8585 | 1.22 | |
| BaselineSchedule=Standard2025.12 | 0.858 | 1 | |
| Exp-k=16 (7.5, 0)Schedule Type=Exponential, k=16, Thresholds (τ_high, τ_low)=(7.5, 0)2025.12 | 0.8455 | 5 | |
| Exp-k=16 (7.5, 2.5)Schedule Type=Exponential, k=16, Thresholds (τ_high, τ_low)=(7.5, 2.5)2025.12 | 0.8449 | 2.29 | |
| Exp-k=8 (7.5, 0)Backbone=LLaDA Base, Curvature (k)=82025.12 | 0.7869 | 3.99 | |
| BaselineBackbone=LLaDA Base2025.12 | 0.7798 | 1 | |
| BaselineFormat=Baseline, Bit width (b)=16.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 0.737 | — | |
| Tensor RMS + CFormat=Tensor RMS + C, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 0.728 | — | |
| Block AbsmaxFormat=Block Absmax, Bit width (b)=3.25, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 0.727 | — | |
| ProphetSchedule=Prophet2025.12 | 0.7258 | 1.03 | |
| Tensor RMS + SpFormat=Tensor RMS + Sp, Bit width (b)=3.05, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 0.716 | — | |
| Channel AbsmaxFormat=Channel Absmax, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 0.693 | — | |
| Tensor AbsmaxFormat=Tensor Absmax, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 0.665 | — | |
| Tensor RMSFormat=Tensor RMS, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 0.514 | — |