Reasoning on HellaSwag
91.84HellaSwag AccuracyMistral-Small
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Mistral-SmallModel Size=24B, GPU Hours=∼1.6M2025.05 | 91.84 | — | — | — | |
| Llama-3.3-70B-InstructShots=102026.04 | 89.8 | — | — | — | |
| Qwen3.5-9BShots=102026.04 | 88.51 | — | — | — | |
| Qwen2.5-InstructModel Size=14B, GPU Hours=∼1.8M2025.05 | 88.28 | — | — | — | |
| InfiFusion*Model Size=14B, GPU Hours=160, Subset Source Models=true2025.05 | 87.91 | — | — | — | |
| FuseLLM*Model Size=14B, GPU Hours=225, Subset Source Models=true2025.05 | 87.81 | — | — | — | |
| SFTModel Size=14B, GPU Hours=152025.05 | 87.75 | — | — | — | |
| Phi-4Model Size=14B, GPU Hours=∼1.0M2025.05 | 87.62 | — | — | — | |
| FuseChat*Model Size=14B, GPU Hours=650, Subset Source Models=true2025.05 | 87.42 | — | — | — | |
| SFT-IPOModel Size=14B, GPU Hours=502025.05 | 87.36 | — | — | — | |
| InfiFPO*Model Size=14B, GPU Hours=55, Subset Source Models=true2025.05 | 87.36 | — | — | — | |
| SFT-DPOModel Size=14B, GPU Hours=502025.05 | 87.31 | — | — | — | |
| SFT-WRPOModel Size=14B, GPU Hours=572025.05 | 87.3 | — | — | — | |
| InfiFPOModel Size=14B, GPU Hours=582025.05 | 87.29 | — | — | — | |
| Qwen3-14BShots=102026.04 | 86.7 | — | — | — | |
| Qwen3-30B-A3B-Inst-25072026.02 | 86.31 | — | 1 | — | |
| InternLM2-Chat-20B-SFTEvaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~20B, Fine-tuning Strategy=SFT2024.03 | 85.9 | — | — | — | |
| InternLM2-Chat-20BEvaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~20B2024.03 | 85.8 | — | — | — | |
| LLaDA2.1-flashInference Mode=S Mode2026.02 | 85.6 | — | 2.31 | — | |
| LLaDA2.1-flashInference Mode=Q Mode2026.02 | 85.31 | — | 1.51 | — | |
| LLaDA2.0-flash2026.02 | 84.97 | — | 1.26 | — | |
| InternLM2-Chat-7B-SFTEvaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~7B, Fine-tuning Strategy=SFT2024.03 | 83.5 | — | — | — | |
| Gemma-3-InstructModel Size=12B2025.05 | 83.34 | — | — | — | |
| InternLM2-Chat-7BEvaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~7B2024.03 | 83 | — | — | — | |
| SecGPT-14BShots=102026.04 | 82.69 | — | — | — | |
| Mixtral-8x7B-v0.1Evaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~20B2024.03 | 81.9 | — | — | — | |
| InternLM2-20BEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~20B2024.03 | 81.6 | — | — | — | |
| Ling-flash-2.02026.02 | 81.59 | — | 1 | — | |
| Qwen3-8BShots=102026.04 | 81.33 | — | — | — | |
| Llama-3.1-8B-InstructShots=102026.04 | 80.9 | — | — | — | |
| Qwen-14BEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~20B2024.03 | 80.3 | — | — | — | |
| Mixtral-8x7B-Instruct-v0.1Evaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~20B2024.03 | 80.3 | — | — | — | |
| Qwen2.5-CoderModel Size=14B, GPU Hours=∼1.8M2025.05 | 79.83 | — | — | — | |
| Qwen3-8Bno think=true2026.02 | 79.56 | — | — | — | |
| InternLM2-7BEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~7B2024.03 | 79.3 | — | — | — | |
| Qwen-14B-ChatEvaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~20B2024.03 | 79.2 | — | — | — | |
| LLAMA-3-8BModel=LLAMA-3-8B, Bits=FP16, Zero-shot=true2026.04 | 79.13 | — | — | — | |
| Llama 3.1 8BAttn CR=N/A2025.08 | 79.1 | — | — | — | |
| LLaDA2.0-mini2026.02 | 79.01 | — | 1.5 | — | |
| Mistral-7B-v0.1Evaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~7B2024.03 | 78.9 | — | — | — | |
| T2MBackbone=LLaDA2.1-mini, Inference Strategy=Token-to-Mask, Remasking Strategy=LOWPROB, τ=0.3, Cmax=1, ρmax=0.252026.04 | 78.67 | — | — | — | |
| Original (T2T)Backbone=LLaDA2.1-mini, Inference Strategy=Text-to-Text2026.04 | 78.57 | — | — | — | |
| InternLM2-20B-BaseEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~20B2024.03 | 78.1 | — | — | — | |
| Matrix PCABase Model=Llama 3.1 8B, Attn CR=20%2025.08 | 78 | — | — | — | |
| Llama2-13BEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~20B2024.03 | 77.5 | — | — | — | |
| SVD-LLMBase Model=Llama 3.1 8B, Attn CR=20%2025.08 | 77.5 | — | — | — | |
| XekRung-8BShots=102026.04 | 77.43 | — | — | — | |
| ChatGLM3-6B-BaseEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~7B2024.03 | 76.5 | — | — | — | |
| InternLM2-7B-BaseEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~7B2024.03 | 76.3 | — | — | — | |
| LLaDA2.1-minimode=Q Mode2026.02 | 76.19 | — | 1.49 | — | |
| LLaDA2.1-minimode=S Mode2026.02 | 75.71 | — | 2.39 | — | |
| Qwen-7BEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~7B2024.03 | 75 | — | — | — | |
| DenseLLM=Llama2-13B, Ratio=0.00%2024.03 | 74.78 | — | — | — | |
| GenKnowSubSetting=Fr, Zero-shot=true, Base Model=Phi-32025.05 | 74.02 | — | — | — | |
| Llama2-7BEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~7B2024.03 | 74 | — | — | — | |
| Baichuan2-13B-BaseEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~20B2024.03 | 73.7 | — | — | — | |
| FP16Model=LLaMA-3.2-3B, Compression Ratio=1x2026.01 | 73.65 | — | — | — | |
| Llama 3.2 3BAttn CR=N/A2025.08 | 73.6 | — | — | — | |
| Phi-3Zero-shot=true, Base Model=Phi-32025.05 | 73.59 | — | — | — | |
| GenKnowSubSetting=Avg, Zero-shot=true, Base Model=Phi-32025.05 | 73.45 | — | — | — | |
| GenKnowSubSetting=En, Zero-shot=true, Base Model=Phi-32025.05 | 73.36 | — | — | — | |
| ChatGLM3-6BEvaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~7B2024.03 | 73.1 | — | — | — | |
| LLAMA-1-7BModel=LLAMA-1-7B, Bits=FP16, Zero-shot=true2026.04 | 73 | — | — | — | |
| LLAMA-2-7BModel=LLAMA-2-7B, Bits=FP16, Zero-shot=true2026.04 | 72.96 | — | — | — | |
| GenKnowSubSetting=De, Zero-shot=true, Base Model=Phi-32025.05 | 72.79 | — | — | — | |
| Mean NormZero-shot=true, Base Model=Phi-32025.05 | 72.53 | — | — | — | |
| SharedZero-shot=true, Base Model=Phi-32025.05 | 72.16 | — | — | — | |
| ArrowZero-shot=true, Base Model=Phi-32025.05 | 71.89 | — | — | — | |
| QMCModel=LLaMA-3.2-3B, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 71.77 | — | — | — | |
| Matrix PCABase Model=Llama 3.2 3B, Attn CR=20%2025.08 | 71.3 | — | — | — | |
| DenseLLM=Llama2-7B, Ratio=0.00%2024.03 | 71.26 | — | — | — | |
| FP16Model=Hymba-Instruct-1.5B, Compression Ratio=1x2026.01 | 71.1 | — | — | — | |
| SVD-LLMBase Model=Llama 3.2 3B, Attn CR=20%2025.08 | 70.5 | — | — | — | |
| Baichuan2-7B-BaseEvaluation Protocol=0-shot, Model Category=Base Model, Model Scale=~7B2024.03 | 70.2 | — | — | — | |
| GPT-3.5Evaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=API Models2024.03 | 70.2 | — | — | — | |
| QMCModel=Hymba-Instruct-1.5B, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 70.06 | — | — | — | |
| QMCModel=Hymba-Instruct-1.5B, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 69.35 | — | — | — | |
| Ling-mini-2.02026.02 | 69.02 | — | — | — | |
| SmolLM 2Params.=1.7B, Tokens=11T2025.06 | 68.7 | — | — | — | |
| QMCModel=LLaMA-3.2-3B, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 68.4 | — | — | — | |
| FP16Model=Qwen2.5-1.5B-Instruct, Compression Ratio=1x2026.01 | 68.25 | — | — | — | |
| LLMPrun.LLM=Llama2-13B, Ratio=24.4%2024.03 | 67.76 | — | — | — | |
| DenseLLM=Baichuan2-7B, Ratio=0.00%2024.03 | 67.56 | — | — | — | |
| Stable LM 2Params.=1.6B, Tokens=2T2025.06 | 66.7 | — | — | — | |
| ShortGPTLLM=Llama2-13B, Ratio=24.6%2024.03 | 66.64 | — | — | — | |
| Qwen 2.5Params.=1.5B2025.06 | 66.5 | — | — | — | |
| QMCModel=Qwen2.5-1.5B-Instruct, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 66.42 | — | — | — | |
| LaCoLLM=Llama2-13B, Ratio=24.6%2024.03 | 64.39 | — | — | — | |
| Mistral-7B-Instruct-v0.2Evaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~7B2024.03 | 64.1 | — | — | — | |
| AdamParams.=1.4B, Tokens=1T2025.06 | 64 | — | — | — | |
| Qwen 2Params.=1.5B, Tokens=7T2025.06 | 63.9 | — | — | — | |
| Llama 3.2 1BAttn CR=N/A2025.08 | 63.7 | — | — | — | |
| QMCModel=Qwen2.5-1.5B-Instruct, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 63.49 | — | — | — | |
| MXINT4Model=Hymba-Instruct-1.5B, Compression Ratio=4x2026.01 | 63.33 | — | — | — | |
| SmolLMParams.=1.7B, Tokens=1T2025.06 | 63 | — | — | — | |
| FP16Model=Phi-1.5B, Compression Ratio=1x2026.01 | 62.68 | — | — | — | |
| Qwen-7B-ChatEvaluation Protocol=0-shot, Model Category=Chat Model, Model Scale=~7B2024.03 | 61.9 | — | — | — | |
| QMCModel=Phi-1.5B, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 61.79 | — | — | — | |
| MXINT4Model=Qwen2.5-1.5B-Instruct, Compression Ratio=4x2026.01 | 61.39 | — | — | — | |
| LLAMA 3.2Params.=1.2B2025.06 | 61.3 | — | — | — |