Commonsense Question Answering on CSQA
82.72AccuracySAC Single-task
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SAC Single-taskBackbone=Qwen2-72B-Instruct2024.11 | 82.72 | — | |
| SAC Multi-taskBackbone=Qwen2-72B-Instruct2024.11 | 82.56 | — | |
| No ControlBackbone=Qwen2-72B-Instruct2024.11 | 82.39 | — | |
| InfLLMv2Model Size=8B, Context=short-context2026.05 | 80.6 | — | |
| NSAModel Size=8B, Context=short-context2026.05 | 80.4 | — | |
| DashAttentionModel Size=8B, Context=short-context2026.05 | 80.3 | — | |
| FullAttnModel Size=8B, Context=short-context2026.05 | 80 | — | |
| No ControlBackbone=Qwen2-7B-Instruct2024.11 | 78.62 | — | |
| SAC Single-taskBackbone=Qwen2-7B-Instruct2024.11 | 77.07 | — | |
| SAC Multi-taskBackbone=Qwen2-7B-Instruct2024.11 | 76.9 | — | |
| Top-K ReducedSparsity Level=Mid Sparsity, K=4, Avg. K=4.002026.05 | 72.48 | — | |
| MoE-DynamicSparsity Level=High Sparsity, phi=0.1, Avg. K=3.902026.05 | 71.09 | — | |
| BEAMSparsity Level=High Sparsity, beta=0.1, Avg. K=1.082026.05 | 70.93 | — | |
| MoE-DynamicSparsity Level=Mid Sparsity, phi=0.3, Avg. K=4.312026.05 | 70.6 | — | |
| AdaMoESparsity Level=High Sparsity, Null=128, Avg. K=2.112026.05 | 70.27 | — | |
| AdaMoESparsity Level=Mid Sparsity, Null=64, Avg. K=3.252026.05 | 69.86 | — | |
| 16-bit BaselineModel Variant=LLaMA-3-8B2024.11 | 69.2 | — | |
| DeepSeekV2-LiteK=6, Avg. K=6.002026.05 | 68.14 | — | |
| BEAMSparsity Level=Extreme Sparsity, beta=1.0, Avg. K=0.482026.05 | 67.57 | — | |
| Top-K PruningSparsity Level=Mid Sparsity, K=4, Avg. K=4.002026.05 | 67.4 | — | |
| 16-bit BaselineModel Variant=LLaMA-2-13B2024.11 | 67.3 | — | |
| Top-K ReducedSparsity Level=High Sparsity, K=2, Avg. K=2.002026.05 | 66.83 | — | |
| ANVFP4Group Size (GS)=16, Model Variant=LLaMA-2-13B2024.11 | 66.2 | — | |
| NVFP4Group Size (GS)=16, Model Variant=LLaMA-2-13B2024.11 | 65.3 | — | |
| BEAMSparsity Level=Mid Sparsity, beta=0.01, Avg. K=2.612026.05 | 65.11 | — | |
| MXFP4Group Size (GS)=32, Model Variant=LLaMA-2-13B2024.11 | 65.1 | — | |
| NVFP4Group Size (GS)=32, Model Variant=LLaMA-2-13B2024.11 | 65 | — | |
| 16-bit BaselineModel Variant=LLaMA-2-7B2024.11 | 64.9 | — | |
| AMXFP4Group Size (GS)=32, Model Variant=LLaMA-2-13B2024.11 | 64.9 | — | |
| ANVFP4Group Size (GS)=16, Model Variant=LLaMA-3-8B2024.11 | 64.9 | — | |
| ANVFP4Group Size (GS)=32, Model Variant=LLaMA-2-13B2024.11 | 64.7 | — | |
| NVFP4Group Size (GS)=16, Model Variant=LLaMA-3-8B2024.11 | 63.4 | — | |
| ANVFP4Group Size (GS)=32, Model Variant=LLaMA-3-8B2024.11 | 62.9 | — | |
| NVFP4Group Size (GS)=16, Model Variant=LLaMA-2-7B2024.11 | 62.6 | — | |
| MXFP4-POTGroup Size (GS)=32, Model Variant=LLaMA-2-13B2024.11 | 62.2 | — | |
| AMXFP4Group Size (GS)=32, Model Variant=LLaMA-3-8B2024.11 | 62.2 | — | |
| ANVFP4Group Size (GS)=32, Model Variant=LLaMA-2-7B2024.11 | 62.2 | — | |
| ANVFP4Group Size (GS)=16, Model Variant=LLaMA-2-7B2024.11 | 62.2 | — | |
| MXFP4Group Size (GS)=32, Model Variant=LLaMA-3-8B2024.11 | 62 | — | |
| AMXFP4Group Size (GS)=32, Model Variant=LLaMA-2-7B2024.11 | 62 | — | |
| NVFP4Group Size (GS)=32, Model Variant=LLaMA-3-8B2024.11 | 61.9 | — | |
| MXFP4Group Size (GS)=32, Model Variant=LLaMA-2-7B2024.11 | 61.6 | — | |
| NVFP4Group Size (GS)=32, Model Variant=LLaMA-2-7B2024.11 | 61.4 | — | |
| MXFP4-POTGroup Size (GS)=32, Model Variant=LLaMA-2-7B2024.11 | 59.4 | — | |
| Top-K PruningSparsity Level=High Sparsity, K=2, Avg. K=2.002026.05 | 58.64 | — | |
| MXFP4-POTGroup Size (GS)=32, Model Variant=LLaMA-3-8B2024.11 | 58.6 | — | |
| Top-K ReducedSparsity Level=Extreme Sparsity, K=1, Avg. K=1.002026.05 | 52.91 | — | |
| MIDUS-HMLTraining Stage=SFT-1B2025.12 | 50.04 | — | |
| OpT-DeUSTraining Stage=SFT-1B2025.12 | 49.55 | — | |
| LESATraining Stage=SFT-1B2025.12 | 47.91 | — | |
| Avg-DeUSTraining Stage=SFT-1B2025.12 | 47.01 | — | |
| OpT-DeUSTraining Stage=CPT-1B2025.12 | 46.44 | — | |
| MIDUS-HMLTraining Stage=CPT-1B2025.12 | 46.27 | — | |
| LESATraining Stage=CPT-1B2025.12 | 45.21 | — | |
| Llama ProTraining Stage=SFT-1B2025.12 | 42.59 | — | |
| Avg-DeUSTraining Stage=CPT-1B2025.12 | 41.93 | — | |
| Llama ProTraining Stage=CPT-1B2025.12 | 35.14 | — | |
| SOLARTraining Stage=SFT-1B2025.12 | 27.44 | — | |
| BaseTraining Stage=SFT-1B2025.12 | 26.37 | — | |
| SOLARTraining Stage=CPT-1B2025.12 | 26.29 | — | |
| BaseTraining Stage=CPT-1B2025.12 | 25.88 | — | |
| Gemma 7BFramework=Adaptive HPF2024.06 | — | 5.5771 | |
| Llama-3 8BFramework=Adaptive HPF2024.06 | — | 5.9136 | |
| Mistral 7BFramework=Adaptive HPF2024.06 | — | 5.9073 | |
| Phi-3 3.8BFramework=Adaptive HPF2024.06 | — | 5.6793 |