Commonsense Reasoning on HellaSwag (HS) (Accuracy)
78.94HS AccuracyOriginal
Evaluation Results
| Method | Links | |
|---|---|---|
| OriginalBase Model=LLaMA3-8B2026.05 | 78.94 | |
| OriginalBackbone=LLaMA3-8B2026.05 | 78.94 | |
| Accuracy (Arc-E)Backbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 73.62 | |
| LLaMA3Size=3.0B, Bit=16.0, Zero-shot=true2025.08 | 73.6 | |
| Accuracy (Ours)Backbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 73.41 | |
| Llama-3.1-8B-InstructModel=Llama-3.1-8B-Instruct, Sparsity=Dense2026.06 | 72.52 | |
| DeepSeek-7B-chatModel=DeepSeek-7B-chat, Sparsity=Dense2026.06 | 70.56 | |
| Cosine SimilarityPruning Method=Cosine Similarity, Base Model=LLaMA3-8B, Calibration Dataset Size=1,500-instance2026.05 | 70.24 | |
| Cosine SimilarityBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 69.35 | |
| Out. Norm-SimPruning Method=Out. Norm-Sim, Base Model=LLaMA3-8B, Calibration Dataset Size=1,500-instance2026.05 | 67 | |
| Out. Cosine-SimPruning Method=Out. Cosine-Sim, Base Model=LLaMA3-8B, Calibration Dataset Size=1,500-instance2026.05 | 66.57 | |
| PerplexityPruning Method=Perplexity, Base Model=LLaMA3-8B, Calibration Dataset Size=1,500-instance2026.05 | 65.96 | |
| Qwen2.5-7B-InstructModel=Qwen2.5-7B-Instruct, Sparsity=Dense2026.06 | 65.48 | |
| Out. Divergence-SimPruning Method=Out. Divergence-Sim, Base Model=LLaMA3-8B, Calibration Dataset Size=1,500-instance2026.05 | 65.09 | |
| Out. Cosine-SimBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 64.6 | |
| Llama-3.2-3B-InstructModel=Llama-3.2-3B-Instruct, Sparsity=Dense2026.06 | 64.53 | |
| Out. Divergence-SimBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 64.37 | |
| TaylorPruning Method=Taylor, Base Model=LLaMA3-8B, Calibration Dataset Size=1,500-instance2026.05 | 63.73 | |
| LLaMA3Size=1.3B, Bit=16.0, Zero-shot=true2025.08 | 63.7 | |
| TaylorBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 62.73 | |
| Out. Norm-SimBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 61.73 | |
| Slice-GPTBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 61.19 | |
| OPTSize=2.7B, Bit=16.0, Zero-shot=true2025.08 | 60.6 | |
| ReplaceMe (Cosine)Model=Llama-3.1-8B-Instruct, Sparsity=25%2026.06 | 60.22 | |
| Acc (Ours)Pruning Method=Accuracy-based relevance score, Base Model=LLaMA3-8B, Calibration Dataset Size=1,500-instance2026.05 | 60.16 | |
| Accuracy (C4)Backbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 60.16 | |
| Streamline (Layer)Model=Llama-3.1-8B-Instruct, Sparsity=25%2026.06 | 59.9 | |
| SUBFITModel=Llama-3.1-8B-Instruct, Sparsity=25%2026.06 | 59.33 | |
| PythiaSize=2.8B, Bit=16.0, Zero-shot=true2025.08 | 59.3 | |
| Streamline (FFN)Model=Llama-3.1-8B-Instruct, Sparsity=25%2026.06 | 58.31 | |
| BinaryLLMSize=3.0B, Bit=1.01, Zero-shot=true2025.08 | 58.3 | |
| ReplaceMe (LS)Model=Llama-3.1-8B-Instruct, Sparsity=25%2026.06 | 57.27 | |
| OPTSize=1.3B, Bit=16.0, Zero-shot=true2025.08 | 53.7 | |
| ReplaceMe (Cosine)Model=Qwen2.5-7B-Instruct, Sparsity=25%2026.06 | 53.17 | |
| Streamline (Layer)Model=Qwen2.5-7B-Instruct, Sparsity=25%2026.06 | 53.1 | |
| ReplaceMe (LS)Model=Qwen2.5-7B-Instruct, Sparsity=25%2026.06 | 52.47 | |
| Streamline (FFN)Model=Qwen2.5-7B-Instruct, Sparsity=25%2026.06 | 51.87 | |
| SUBFITModel=Llama-3.2-3B-Instruct, Sparsity=25%2026.06 | 51.42 | |
| Streamline (Layer)Model=Llama-3.2-3B-Instruct, Sparsity=25%2026.06 | 50.93 | |
| ReplaceMe (Cosine)Model=DeepSeek-7B-chat, Sparsity=25%2026.06 | 50.7 | |
| ReplaceMe (Cosine)Model=Llama-3.2-3B-Instruct, Sparsity=25%2026.06 | 50.45 | |
| Streamline (FFN)Model=Llama-3.2-3B-Instruct, Sparsity=25%2026.06 | 50.25 | |
| Streamline (Layer)Model=DeepSeek-7B-chat, Sparsity=25%2026.06 | 50.02 | |
| BinaryLLMSize=1.3B, Bit=1.01, Zero-shot=true2025.08 | 49.6 | |
| ReplaceMe (LS)Model=Llama-3.2-3B-Instruct, Sparsity=25%2026.06 | 48.84 | |
| SUBFITModel=DeepSeek-7B-chat, Sparsity=25%2026.06 | 48.64 | |
| SUBFITModel=Qwen2.5-7B-Instruct, Sparsity=25%2026.06 | 48.19 | |
| PerplexityBackbone=LLaMA3-8B, Pruning Ratio=25%2026.05 | 48.01 | |
| ReplaceMe (LS)Model=DeepSeek-7B-chat, Sparsity=25%2026.06 | 47.15 | |
| Qwen3-4B-InstructModel=Qwen3-4B-Instruct, Sparsity=Dense2026.06 | 43.42 | |
| BitNetSize=3.0B, Bit=1.59, Zero-shot=true2025.08 | 42.9 | |
| SmolLMSize=135M, Bit=16.0, Zero-shot=true2025.08 | 42.6 | |
| FBI-LLMSize=1.3B, Bit=1.01, Zero-shot=true2025.08 | 42.3 | |
| Streamline (FFN)Model=DeepSeek-7B-chat, Sparsity=25%2026.06 | 39.83 | |
| OneBit-OPTSize=2.7B, Bit=1.02, Zero-shot=true2025.08 | 38.2 | |
| BitNetSize=1.3B, Bit=1.59, Zero-shot=true2025.08 | 37.7 | |
| ReplaceMe (LS)Model=Qwen3-4B-Instruct, Sparsity=25%2026.06 | 37.05 | |
| ReplaceMe (Cosine)Model=Qwen3-4B-Instruct, Sparsity=25%2026.06 | 36.72 | |
| OneBit-OPTSize=1.3B, Bit=1.02, Zero-shot=true2025.08 | 34.3 | |
| OPTSize=125M, Bit=16.0, Zero-shot=true2025.08 | 31.3 | |
| PythiaSize=160M, Bit=16.0, Zero-shot=true2025.08 | 31.3 | |
| BinaryLLMSize=135M, Bit=1.01, Zero-shot=true2025.08 | 31.2 | |
| SUBFITModel=Qwen3-4B-Instruct, Sparsity=25%2026.06 | 28.88 | |
| FBI-LLMSize=130M, Bit=1.01, Zero-shot=true2025.08 | 28.7 | |
| Streamline (FFN)Model=Qwen3-4B-Instruct, Sparsity=25%2026.06 | 22.93 | |
| Streamline (Layer)Model=Qwen3-4B-Instruct, Sparsity=25%2026.06 | 22.63 |