Multiple Choice Question Answering on HellaSwag
93.59AccuracyStable-LoRA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Stable-LoRAModel=3B2026.03 | 93.59 | — | |
| LoRA+Model=3B2026.03 | 93.54 | — | |
| AdamWModel=3B2026.03 | 93.39 | — | |
| LoRA-RITEModel=3B2026.03 | 93.21 | — | |
| RiemannModel=3B2026.03 | 92.61 | — | |
| Stable-LoRAModel=1.5B2026.03 | 88.52 | — | |
| AdamWModel=1.5B2026.03 | 88.28 | — | |
| LoRA-RITEModel=1.5B2026.03 | 87.97 | — | |
| LoRA+Model=1.5B2026.03 | 87.94 | — | |
| RiemannModel=1.5B2026.03 | 86.46 | — | |
| Stable-LoRAModel=1B2026.03 | 84.41 | — | |
| MA-PoP2026.05 | 84.33 | — | |
| Decentr. MADT=52026.05 | 84 | — | |
| AdamWModel=1B2026.03 | 83.76 | — | |
| LoRA+Model=1B2026.03 | 83.39 | — | |
| Decentr. MADT=32026.05 | 83.33 | — | |
| Gemma-9B2026.05 | 82.67 | — | |
| Decentr. MADT=22026.05 | 82.67 | — | |
| ISP2026.05 | 82.67 | — | |
| LoRA-RITEModel=1B2026.03 | 82.38 | — | |
| Self-Consistency2026.05 | 82.33 | — | |
| Sparse MADT=22026.05 | 82 | — | |
| Sparse MADT=32026.05 | 81.33 | — | |
| Sparse MADT=52026.05 | 81.33 | — | |
| Majority Voting2026.05 | 80.33 | — | |
| Free MADT=32026.05 | 79.67 | — | |
| Free MADT=52026.05 | 79.33 | — | |
| FP16Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 79.2 | — | |
| LLaMA-3 8BPruning Ratio=0%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 79.19 | — | |
| Free MADT=22026.05 | 79 | — | |
| Qwen-7B2026.05 | 78.8 | — | |
| LLP2026.05 | 78.33 | — | |
| K-stage RKSolver choice=Runge–Kutta, K=22026.05 | 77.93 | — | |
| Heun K=1Solver choice=Heun, K=12026.05 | 77.78 | — | |
| Baseline2026.05 | 77.74 | — | |
| RiemannModel=1B2026.03 | 77.41 | — | |
| DuQuant++Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 77.3 | — | |
| FP16Backbone=LLaMA-3-8B, Evaluation Protocol=zero-shot2026.06 | 77.13 | — | |
| DuQuant++*Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 77 | — | |
| Centr. MADT=22026.05 | 76.67 | — | |
| Centr. MADT=52026.05 | 76.33 | — | |
| MR-GPTQModel=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 76.2 | — | |
| LLaMA-2 7BPruning Ratio=0%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 75.99 | — | |
| FP16Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 75.7 | — | |
| FlatQuantModel=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 75.7 | — | |
| Centr. MADT=32026.05 | 75.67 | — | |
| FP16Backbone=Qwen-3-8B, Evaluation Protocol=zero-shot2026.06 | 74.38 | — | |
| DuQuant++Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 73.8 | — | |
| DuQuant++*Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 73.7 | — | |
| MR-GPTQModel=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 73.2 | — | |
| SAGE-PTQBackbone=LLaMA-3-8B, Evaluation Protocol=zero-shot2026.06 | 73.15 | — | |
| FlatQuantModel=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 73 | — | |
| FP16Bits=-, Backbone=LLAMA-2-7B, Zero-shot=true2026.04 | 72.96 | — | |
| Dense baselineBase Model=Llama-3.1-8B-Instruct, Sparsity=Dense2026.06 | 72.52 | — | |
| QuaRotModel=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 72.4 | — | |
| MoPPruning Ratio=20%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 71.94 | — | |
| QuaRot*Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 71.8 | — | |
| SAGE-PTQBackbone=Qwen-3-8B, Evaluation Protocol=zero-shot2026.06 | 71.61 | — | |
| Falcon-7B2026.05 | 71.33 | — | |
| MoPPruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 71.16 | — | |
| SlimLLMPruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 70.95 | — | |
| AmoebaLLMPruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 70.8 | — | |
| Dense baselineBase Model=DeepSeek-7B-chat, Sparsity=Dense2026.06 | 70.56 | — | |
| Pretrained ModelModel Size=7B, Training Strategy=Standard Pretraining2026.03 | 69.39 | — | |
| Split Model TrainingModel Size=2.7B, Training Strategy=Split Training2026.03 | 69.37 | — | |
| LINEARPATCHPruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 69.33 | — | |
| AMPPruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 69.22 | — | |
| QuaRot*Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 69.2 | — | |
| CoMePruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 68.68 | — | |
| PruneNetPruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 68.37 | — | |
| SUBFITBase Model=Llama-3.1-8B-Instruct, Sparsity=12.5%2026.06 | 68.19 | — | |
| QuaRotModel=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 68 | — | |
| Pretrained ModelModel Size=2.7B, Training Strategy=Standard Pretraining2026.03 | 67.6 | — | |
| Split Model TrainingModel Size=1.3B, Training Strategy=Split Training2026.03 | 67.3 | — | |
| MoPPruning Ratio=30%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 66.88 | — | |
| LINEARPATCHPruning Ratio=20%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 66.74 | — | |
| Stable-LoRAModel=0.5B2026.03 | 66.73 | — | |
| AMPPruning Ratio=20%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 66.53 | — | |
| CoMePruning Ratio=30%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 65.83 | — | |
| AdamWModel=0.5B2026.03 | 65.65 | — | |
| CoMePruning Ratio=20%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 65.52 | — | |
| Dense baselineBase Model=Qwen2.5-7B-Instruct, Sparsity=Dense2026.06 | 65.48 | — | |
| AMPPruning Ratio=30%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 65.47 | — | |
| MoPPruning Ratio=30%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 65.44 | — | |
| LoRA+Model=0.5B2026.03 | 64.75 | — | |
| Dense baselineBase Model=Llama-3.2-3B-Instruct, Sparsity=Dense2026.06 | 64.53 | — | |
| LINEARPATCHPruning Ratio=30%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 64.52 | — | |
| Mistral-7B2026.05 | 64.33 | — | |
| FP16Backbone=OPT-6.7B, Evaluation Protocol=zero-shot2026.06 | 64.09 | — | |
| Pretrained ModelModel Size=1.3B, Training Strategy=Standard Pretraining2026.03 | 63.91 | — | |
| PruneNetPruning Ratio=30%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 63.21 | — | |
| SUBFITBase Model=DeepSeek-7B-chat, Sparsity=12.5%2026.06 | 62.89 | — | |
| DISP-LLMPruning Ratio=30%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 62.87 | — | |
| LoRA-RITEModel=0.5B2026.03 | 62.73 | — | |
| ModHiFi-PPruning Ratio=20%, Base Model=LLaMA-2 7B, Evaluation Protocol=Zero-shot2026.02 | 62.7 | — | |
| Llama-8B2026.05 | 62.67 | — | |
| AMPPruning Ratio=30%, Base Model=LLaMA-3 8B, Evaluation Protocol=Zero-shot2026.02 | 61.93 | — | |
| RiemannModel=0.5B2026.03 | 60.79 | — | |
| Streamline (Layer)Base Model=Llama-3.2-3B-Instruct, Sparsity=12.5%2026.06 | 60.16 | — | |
| FP16 (no compression)error budget (epsilon)=N/A, KV cache (rel.)=1, Zero-shot protocol=true, Base Model=Llama-3-8B2026.01 | 60 | — |