Common Sense Reasoning on HellaSwag (test)
83.9AccuracyLlama-3.3-70B-Base
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-3.3-70B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 83.9 | |
| ReM-MoA*Width (N)=8, Depth (L)=3, Trend=↑↑, Distilled Reviewer=true2026.06 | 83.7 | |
| MISTRAL-7BModel Architecture=MISTRAL-7B, Configuration Type=Dense, Evaluation Protocol=Zero-shot, Sparsity Level=0%2026.01 | 83.43 | |
| ReM-MoA*Width (N)=7, Depth (L)=3, Trend=↑↑, Distilled Reviewer=true2026.06 | 83.4 | |
| ReM-MoA*Width (N)=6, Depth (L)=3, Trend=↑↑, Distilled Reviewer=true2026.06 | 83.1 | |
| ReM-MoAWidth (N)=8, Depth (L)=3, Trend=↑↑2026.06 | 82.9 | |
| ReM-MoAWidth (N)=7, Depth (L)=3, Trend=↑↑2026.06 | 82.6 | |
| ReM-MoA*Width (N)=5, Depth (L)=3, Trend=↑↑, Distilled Reviewer=true2026.06 | 82.6 | |
| ReM-MoAWidth (N)=6, Depth (L)=3, Trend=↑↑2026.06 | 82.3 | |
| ReM-MoA*Width (N)=4, Depth (L)=3, Trend=↑↑, Distilled Reviewer=true2026.06 | 82 | |
| ReM-MoAWidth (N)=5, Depth (L)=3, Trend=↑↑2026.06 | 81.8 | |
| QWEN3-14BModel Architecture=QWEN3-14B, Configuration Type=Dense, Evaluation Protocol=Zero-shot, Sparsity Level=0%2026.01 | 81.44 | |
| ReM-MoAWidth (N)=4, Depth (L)=3, Trend=↑↑2026.06 | 81.2 | |
| ReM-MoA*Width (N)=3, Depth (L)=3, Trend=↑↑, Distilled Reviewer=true2026.06 | 81.2 | |
| ReM-MoAWidth (N)=3, Depth (L)=3, Trend=↑↑2026.06 | 80.5 | |
| AttentionMoAWidth (N)=8, Depth (L)=3, Trend=↑→2026.06 | 80.4 | |
| AttentionMoAWidth (N)=7, Depth (L)=3, Trend=↑→2026.06 | 80.3 | |
| AttentionMoAWidth (N)=6, Depth (L)=3, Trend=↑→2026.06 | 80.1 | |
| AttentionMoAWidth (N)=5, Depth (L)=3, Trend=↑→2026.06 | 79.7 | |
| LLAMA-3.1-8BModel Architecture=LLAMA-3.1-8B, Configuration Type=Dense, Evaluation Protocol=Zero-shot, Sparsity Level=0%2026.01 | 79.31 | |
| Mistral-3.2-24B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=European2026.02 | 79.3 | |
| AttentionMoAWidth (N)=4, Depth (L)=3, Trend=↑→2026.06 | 79.2 | |
| RMoAWidth (N)=8, Depth (L)=3, Trend=→2026.06 | 78.9 | |
| RMoAWidth (N)=7, Depth (L)=3, Trend=→2026.06 | 78.8 | |
| RMoAWidth (N)=6, Depth (L)=3, Trend=→2026.06 | 78.7 | |
| FP16Model=Llama-3-8B, Quantization Precision=3-bit2025.12 | 78.5 | |
| AttentionMoAWidth (N)=3, Depth (L)=3, Trend=↑→2026.06 | 78.5 | |
| RMoAWidth (N)=5, Depth (L)=3, Trend=→2026.06 | 78.4 | |
| Gemma-3-27B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 78.2 | |
| RMoAWidth (N)=4, Depth (L)=3, Trend=→2026.06 | 78 | |
| Gemma-3-12B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 77.7 | |
| Standard MoAWidth (N)=6, Depth (L)=3, Trend=→2026.06 | 77.5 | |
| RMoAWidth (N)=3, Depth (L)=3, Trend=→2026.06 | 77.5 | |
| Apertus-70B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 77.4 | |
| Standard MoAWidth (N)=5, Depth (L)=3, Trend=→2026.06 | 77.4 | |
| Standard MoAWidth (N)=7, Depth (L)=3, Trend=→2026.06 | 77.4 | |
| OLMo-3-32B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=Non-European2026.02 | 77.2 | |
| Standard MoAWidth (N)=4, Depth (L)=3, Trend=→2026.06 | 77.2 | |
| Standard MoAWidth (N)=8, Depth (L)=3, Trend=→2026.06 | 77.2 | |
| Standard MoAWidth (N)=3, Depth (L)=3, Trend=→2026.06 | 76.8 | |
| Qwen-3-30B-A3B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 76.5 | |
| Qwen-3-14B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 76.2 | |
| DEEPSEEK-R1-QWEN-8BModel Architecture=DEEPSEEK-R1-QWEN-8B, Configuration Type=Dense, Evaluation Protocol=Zero-shot, Sparsity Level=0%2026.01 | 75.72 | |
| Llama-3.1-8B-Baseshot=3-shot, approach=likelihood-based, access_level=Open-weights, region=Non-European2026.02 | 75.7 | |
| DEEPSEEK-R1-LLAMA-8BModel Architecture=DEEPSEEK-R1-LLAMA-8B, Configuration Type=Dense, Evaluation Protocol=Zero-shot, Sparsity Level=0%2026.01 | 74.78 | |
| CALMModel=Llama-3-8B, Quantization Precision=3-bit2025.12 | 74.5 | |
| LLAMA-3.2-3BModel Architecture=LLAMA-3.2-3B, Configuration Type=Dense, Evaluation Protocol=Zero-shot, Sparsity Level=0%2026.01 | 74.13 | |
| DARTModel Architecture=MISTRAL-7B, Configuration Type=Sparse, Evaluation Protocol=Zero-shot, Sparsity Level=70%2026.01 | 74.12 | |
| QWEN3-4BModel Architecture=QWEN3-4B, Configuration Type=Dense, Evaluation Protocol=Zero-shot, Sparsity Level=0%2026.01 | 73.75 | |
| EuroLLM-9B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 73.6 | |
| EuroLLM-22B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 73.2 | |
| Apertus-8B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=European2026.02 | 73.2 | |
| SpinQuantModel=Llama-3-8B, Quantization Precision=3-bit2025.12 | 72.1 | |
| SmoothQuantModel=Llama-3-8B, Quantization Precision=3-bit2025.12 | 71.8 | |
| FP16Model=Llama-3.2-3B, Quantization Precision=3-bit2025.12 | 70.2 | |
| AWQModel=Llama-3-8B, Quantization Precision=3-bit2025.12 | 69.5 | |
| OLMo-3-7B-Baseshot=3-shot, approach=likelihood-based, access_level=Fully-open, region=Non-European2026.02 | 69.2 | |
| DARTModel Architecture=QWEN3-14B, Configuration Type=Sparse, Evaluation Protocol=Zero-shot, Sparsity Level=70%2026.01 | 68.4 | |
| CALMModel=Llama-3.2-3B, Quantization Precision=3-bit2025.12 | 66.5 | |
| GPTQModel=Llama-3-8B, Quantization Precision=3-bit2025.12 | 65.2 | |
| DARTModel Architecture=LLAMA-3.1-8B, Configuration Type=Sparse, Evaluation Protocol=Zero-shot, Sparsity Level=70%2026.01 | 64.58 | |
| SpinQuantModel=Llama-3.2-3B, Quantization Precision=3-bit2025.12 | 64.1 | |
| SmoothQuantModel=Llama-3.2-3B, Quantization Precision=3-bit2025.12 | 63.5 | |
| AWQModel=Llama-3.2-3B, Quantization Precision=3-bit2025.12 | 60.2 | |
| DARTModel Architecture=DEEPSEEK-R1-QWEN-8B, Configuration Type=Sparse, Evaluation Protocol=Zero-shot, Sparsity Level=70%2026.01 | 59.49 | |
| EvoGMBackbone=Qwen2.5-1.5B2026.05 | 59.4 | |
| DAREBackbone=Qwen2.5-1.5B2026.05 | 59.3 | |
| CMABackbone=Qwen2.5-1.5B2026.05 | 59.2 | |
| TABackbone=Qwen2.5-1.5B2026.05 | 58.8 | |
| PSO-MergingBackbone=Qwen2.5-1.5B2026.05 | 58.7 | |
| Model SwarmBackbone=Qwen2.5-1.5B2026.05 | 58.7 | |
| DARTModel Architecture=DEEPSEEK-R1-LLAMA-8B, Configuration Type=Sparse, Evaluation Protocol=Zero-shot, Sparsity Level=70%2026.01 | 58.63 | |
| BaseBackbone=Qwen2.5-1.5B2026.05 | 58.1 | |
| Model SoupBackbone=Qwen2.5-1.5B2026.05 | 57.8 | |
| Single BestBackbone=Qwen2.5-1.5B2026.05 | 57.2 | |
| GPTQModel=Llama-3.2-3B, Quantization Precision=3-bit2025.12 | 55.8 | |
| DARTModel Architecture=QWEN3-4B, Configuration Type=Sparse, Evaluation Protocol=Zero-shot, Sparsity Level=70%2026.01 | 53.66 | |
| DARTModel Architecture=LLAMA-3.2-3B, Configuration Type=Sparse, Evaluation Protocol=Zero-shot, Sparsity Level=70%2026.01 | 52.77 | |
| Dip-SVDRatio=0.2, Model=LLaMA-13B, Fine-tuning=false, Mixed-rank strategies=true, Zero-shot evaluation=true2026.02 | 49 | |
| MTLBackbone=Qwen2.5-1.5B2026.05 | 48.4 | |
| SAES-SVDRatio=0.2, Model=LLaMA-13B, Fine-tuning=false, Mixed-rank strategies=false, Zero-shot evaluation=true2026.02 | 47.7 | |
| SVD-LLMRatio=0.2, Model=LLaMA-13B, Fine-tuning=true, Mixed-rank strategies=false, Zero-shot evaluation=true2026.02 | 47 | |
| SAES-SVDRatio=0.4, Model=LLaMA-13B, Fine-tuning=false, Mixed-rank strategies=false, Zero-shot evaluation=true2026.02 | 47 | |
| TIESBackbone=Qwen2.5-1.5B2026.05 | 47 | |
| Dip-SVDRatio=0.4, Model=LLaMA-13B, Fine-tuning=false, Mixed-rank strategies=true, Zero-shot evaluation=true2026.02 | 40.2 | |
| SVD-LLMRatio=0.4, Model=LLaMA-13B, Fine-tuning=true, Mixed-rank strategies=false, Zero-shot evaluation=true2026.02 | 35.5 |