Commonsense Reasoning on SIQA
89.85AccuracyIn-Squeeze
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| In-Squeezerank=128 to 1, schedule=Min steps2026.02 | 89.85 | — | — | |
| In-Squeezerank=128 to 1, schedule=Standard2026.02 | 89.77 | — | — | |
| Direct Fine-tuningrank=1, additional_steps=2002026.02 | 89.35 | — | — | |
| Cont-Squeezerank=128 to 1, additional_steps=7002026.02 | 89.35 | — | — | |
| Cont-Squeezerank=128 to 1, additional_steps=2002026.02 | 89.33 | — | — | |
| Direct Fine-tuningrank=1, additional_steps=02026.02 | 88.78 | — | — | |
| Cont-Squeezerank=128 to 1, additional_steps=02026.02 | 88.13 | — | — | |
| Direct Fine-tuningrank=1, additional_steps=7002026.02 | 87.87 | — | — | |
| BaLoRAModel=Llama-3-8B, Params (%)=0.71452026.04 | 81.78 | — | — | |
| FFTModel=Llama-3.2-3B, #Params=3.21B2025.09 | 81.49 | — | — | |
| BoHAModel=Llama-3.2-3B, #Params=48.63M2025.09 | 81.35 | — | — | |
| HiRAModel=Llama-3-8B, Params (%)=0.70022026.04 | 81.15 | — | — | |
| GraLoRAModel=Llama-3.2-3B, #Params=48.63M2025.09 | 81.03 | — | — | |
| ABBAModel=Llama-3.2-3B, #Params=48.63M2025.09 | 81 | — | — | |
| HiRAModel=Llama-3.2-3B, #Params=48.63M2025.09 | 80.81 | — | — | |
| MoRAModel=Llama-3-8B, Params (%)=0.69972026.04 | 80.71 | — | — | |
| BaLoRAModel=Llama-2-7B, Params (%)=0.84272026.04 | 80.45 | — | — | |
| LoRAModel=Llama-3-8B, Params (%)=0.70022026.04 | 79.9 | — | — | |
| DoRAModel=Llama-3-8B, Params (%)=0.70022026.04 | 79.9 | — | — | |
| LoRAModel=Llama-3.2-3B, #Params=48.63M2025.09 | 79.84 | — | — | |
| DoRAModel=Llama-3.2-3B, #Params=49.40M2025.09 | 79.84 | — | — | |
| MoRAModel=Llama-2-7B, Params (%)=0.82412026.04 | 79.53 | — | — | |
| HiRAModel=Llama-2-7B, Params (%)=0.82562026.04 | 79.53 | — | — | |
| LoRAModel=Llama-2-7B, Params (%)=0.82562026.04 | 79.5 | — | — | |
| PiSSAModel=Llama-3.2-3B, #Params=48.63M2025.09 | 79.43 | — | — | |
| rsLoRAModel=Llama-3.2-3B, #Params=48.63M2025.09 | 78.92 | — | — | |
| DoRAModel=Llama-2-7B, Params (%)=0.82562026.04 | 76 | — | — | |
| PiSSAModel=Llama-3.2-1B, #Params=22.54M2025.09 | 75.54 | — | — | |
| FFTModel=Llama-3.2-1B, #Params=1.24B2025.09 | 75.37 | — | — | |
| ABBAModel=Llama-3.2-1B, #Params=22.54M2025.09 | 74.75 | — | — | |
| GraLoRAModel=Llama-3.2-1B, #Params=22.54M2025.09 | 73.95 | — | — | |
| BoHAModel=Llama-3.2-1B, #Params=22.54M2025.09 | 73.83 | — | — | |
| rsLoRAModel=Llama-3.2-1B, #Params=22.54M2025.09 | 73.47 | — | — | |
| LoRAModel=Llama-3.2-1B, #Params=22.54M2025.09 | 73.39 | — | — | |
| HiRAModel=Llama-3.2-1B, #Params=22.54M2025.09 | 73.39 | — | — | |
| DoRAModel=Llama-3.2-1B, #Params=22.92M2025.09 | 73.34 | — | — | |
| RIDERSBackbone=GPT-3.52024.02 | 72.1 | 29 | — | |
| RIDERSModel=GPT-3.52024.02 | 72.1 | 29 | — | |
| RIDERSBackbone=Mistral-7B2024.02 | 70.6 | 24.5 | — | |
| Chain-of-ThoughtBackbone=GPT-3.52024.02 | 70.6 | 30.3 | — | |
| RIDERSModel=Mistral-7B2024.02 | 70.6 | 24.5 | — | |
| CoTModel=GPT-3.52024.02 | 70.6 | 30.3 | — | |
| Chain-of-ThoughtBackbone=Mistral-7B2024.02 | 70.2 | 28.2 | — | |
| CoTModel=Mistral-7B2024.02 | 70.2 | 28.2 | — | |
| SPS OnlyBackbone=Baichuan2-13B, Strategy=Serial-position Swap2024.02 | 69.2 | 27.8 | — | |
| RIDERSBackbone=Baichuan2-13B2024.02 | 69 | 20.5 | — | |
| RD OnlyBackbone=Baichuan2-13B, Strategy=Residual Decoding2024.02 | 68.9 | 25 | — | |
| Few-shot AnswerBackbone=Baichuan2-13B2024.02 | 68.8 | — | — | |
| ChatGPTModel=ChatGPT2026.04 | 68.5 | — | — | |
| Self-ConsistencyBackbone=Baichuan2-13B2024.02 | 67.8 | 33.5 | — | |
| Chain-of-ThoughtBackbone=Baichuan2-13B2024.02 | 66.5 | 37.5 | — | |
| Least-to-MostBackbone=Baichuan2-13B2024.02 | 66.5 | 36.8 | — | |
| Contrasive CoTBackbone=Baichuan2-13B2024.02 | 65 | 37.3 | — | |
| BaselineFormat=Baseline, Bit width (b)=16.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 64.8 | — | — | |
| Self-RefineBackbone=Baichuan2-13B2024.02 | 62.2 | 43.1 | — | |
| Tensor RMS + CFormat=Tensor RMS + C, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 61.5 | — | — | |
| Tensor RMS + SpFormat=Tensor RMS + Sp, Bit width (b)=3.05, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 61.4 | — | — | |
| Block AbsmaxFormat=Block Absmax, Bit width (b)=3.25, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 59.5 | — | — | |
| Channel AbsmaxFormat=Channel Absmax, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 56.5 | — | — | |
| MobileMoE-LActive Parameters=922M, Total Parameters=5.3B, Few-shot count=0-shot2026.05 | 54.3 | — | — | |
| Llama 3 405BParameters=405B2024.07 | 53.7 | — | — | |
| Qwen2.5-1.5B#Params=1.5B/1.5B, #Tokens=18T, Zero-shot=true2026.02 | 53.5 | — | — | |
| MobileMoE-LMem (GB)=2.75, INT4 QAT=true, Zero-shot=true2026.05 | 52.9 | — | — | |
| Llama 3 70BParameters=70B2024.07 | 52.2 | — | — | |
| Qwen3-1.7B#Params=1.7B/1.7B, #Tokens=36T, Zero-shot=true2026.02 | 52.2 | — | — | |
| GemmaSize=8B, Tokens=6T, Context=8K2024.04 | 51.8 | — | — | |
| Gemma 7BParameters=7B2024.07 | 51.8 | — | — | |
| Mixtral 8×22BParameters=8×22B2024.07 | 51.6 | — | — | |
| LLAMA2Size=13B, Tokens=2T, Context=4K2024.04 | 50.3 | — | — | |
| MobileMoE-MActive Parameters=528M, Total Parameters=2.8B, Few-shot count=0-shot2026.05 | 50.2 | — | — | |
| Tensor AbsmaxFormat=Tensor Absmax, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 49.9 | — | — | |
| MEGALODONSize=7B, Tokens=2T, Context=32K2024.04 | 49.6 | — | — | |
| Llama 3 8BParameters=8B2024.07 | 49.5 | — | — | |
| MobileMoE-MMem (GB)=1.48, INT4 QAT=true, Zero-shot=true2026.05 | 49 | — | — | |
| LLaMA-7BModel Category=LLM, Evaluation Mode=Zero-shot2023.09 | 48.9 | — | — | |
| DREAMLLM-7BModel Category=MLLM, Evaluation Mode=Zero-shot2023.09 | 48.8 | — | — | |
| MPTSize=7B, Tokens=1T, Context=4K2024.04 | 48.5 | — | — | |
| LLAMA2Size=7B, Tokens=2T, Context=4K2024.04 | 48.3 | — | — | |
| Mistral 7BParameters=7B2024.07 | 48.2 | — | — | |
| OLMoE-1B-7B#Params=1.3B/7B, #Tokens=5T, Zero-shot=true2026.02 | 47.8 | — | — | |
| Vicuna-7BModel Category=LLM, Evaluation Mode=Zero-shot2023.09 | 47.5 | — | — | |
| SPES-9B#Params=3.1B/9B, #Tokens=400B, Zero-shot=true, initialized from=pretrained dense model2026.02 | 47.5 | — | — | |
| MobileLLM-ProMem (GB)=0.55, INT4 QAT=true, Zero-shot=true2026.05 | 47.4 | — | — | |
| Qwen2.5-0.5B#Params=0.5B/0.5B, #Tokens=18T, Zero-shot=true2026.02 | 47.1 | — | — | |
| MistralSize=7B, Context=16K2024.04 | 47 | — | — | |
| Qwen3-0.6B#Params=0.6B/0.6B, #Tokens=36T, Zero-shot=true2026.02 | 46.9 | — | — | |
| MobileMoE-SActive Parameters=272M, Total Parameters=1.3B, Few-shot count=0-shot2026.05 | 46.8 | — | — | |
| SYNPROModel Size=1.1B, Alpha (organic token ratio)=10%2026.05 | 46.78 | — | — | |
| SmolLM2-1.7B#Params=1.7B/1.7B, #Tokens=11T, Zero-shot=true2026.02 | 46.7 | — | — | |
| Llama-3.1Size=8B, Type=AR2026.01 | 46.7 | — | — | |
| MoE++ 7B#Params=1.2B/7B, #Tokens=1T, Zero-shot=true2026.02 | 45.7 | — | — | |
| MobileMoE-SMem (GB)=0.68, INT4 QAT=true, Zero-shot=true2026.05 | 45.5 | — | — | |
| A3Size=8B, Type=A32026.01 | 45.2 | — | — | |
| Llama3.2-1B#Params=1.1B/1.1B, #Tokens=9T, Zero-shot=true2026.02 | 45 | — | — | |
| SPES-7B#Params=1.6B/7B, #Tokens=500B, Zero-shot=true2026.02 | 44.8 | — | — | |
| MobiLlama-1.3B#Params=1.3B/1.3B, #Tokens=1.3T, Zero-shot=true2026.02 | 44.7 | — | — | |
| Pythia-1.4B#Params=1.4B/1.4B, #Tokens=300B, Zero-shot=true2026.02 | 44.6 | — | — | |
| Pythia-2.8B#Params=2.8B/2.8B, #Tokens=300B, Zero-shot=true2026.02 | 44.5 | — | — | |
| OPT-2.7B#Params=2.7B/2.7B, #Tokens=180B, Zero-shot=true2026.02 | 44.1 | — | — | |
| OPT-1.3B#Params=1.3B/1.3B, #Tokens=180B, Zero-shot=true2026.02 | 43.7 | — | — |