Physical Commonsense Reasoning on PIQA (Accuracy and Normalized Accuracy)
82.54AccuracyBF16
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BF16Model=Mixtral-8x7B, Precision=BF16, Evaluation Protocol=Zero-shot2026.05 | 82.54 | — | |
| OriginalBackbone=Mistral-7B, Pruning Ratio=0%2026.05 | 82.26 | — | |
| SARQC-GBS(Saliency)Model=Mixtral-8x7B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 81.45 | — | |
| AccBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 80.52 | — | |
| GPTQModel=Mixtral-8x7B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 80.36 | — | |
| SARQC-GBS(Identity)Model=Mixtral-8x7B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 79.87 | — | |
| BF16Model=Qwen3-MoE-30B, Precision=BF16, Evaluation Protocol=Zero-shot2026.05 | 79.82 | — | |
| BLT-DNumber of Parameters=3B, Block size=42026.05 | 79.6 | — | |
| FP16Model=LLaMA2-13B, Precision=FP16, Evaluation Protocol=Zero-shot2026.05 | 79.49 | — | |
| BLTNumber of Parameters=3B2026.05 | 79.38 | — | |
| SARQC-GBS(Saliency)Model=Qwen3-MoE-30B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 79 | — | |
| BF16Model=DeepSeek-MoE-16B, Precision=BF16, Evaluation Protocol=Zero-shot2026.05 | 78.73 | — | |
| Out. Cosine-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 78.56 | — | |
| Out. Norm-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 78.56 | — | |
| SARQC-GBS(Identity)Model=Qwen3-MoE-30B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 78.51 | — | |
| FP16Model=LLaMA2-7B, Precision=FP16, Evaluation Protocol=Zero-shot2026.05 | 78.13 | — | |
| BLT-DNumber of Parameters=3B, Block size=82026.05 | 78.02 | — | |
| SARQC-GBS(Saliency)Model=DeepSeek-MoE-16B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 77.91 | — | |
| GPTQModel=DeepSeek-MoE-16B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 77.8 | — | |
| SARQC-GBS(Identity)Model=DeepSeek-MoE-16B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 77.31 | — | |
| GPTQModel=Qwen3-MoE-30B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 77.09 | — | |
| BLT-DNumber of Parameters=3B, Block size=162026.05 | 76.93 | — | |
| SARQC-GBS(Saliency)Model=LLaMA2-7B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 76.33 | — | |
| BLTParameters=1B, Likelihood-based evaluation=true2026.05 | 75.46 | — | |
| SARQC-GBS(Identity)Model=LLaMA2-7B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 75.19 | — | |
| SARQC-GBS(Saliency)Model=LLaMA2-13B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 74.97 | — | |
| GPTAQModel=LLaMA2-7B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 74.54 | — | |
| BLT-DParameters=1B, Block size=4, Likelihood-based evaluation=true2026.05 | 74.48 | — | |
| SARQC-GBS(Identity)Model=LLaMA2-13B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 74.37 | — | |
| Out. Divergence-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 74.32 | — | |
| GPTAQModel=LLaMA2-13B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 74.27 | — | |
| GPTQModel=LLaMA2-7B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 73.94 | — | |
| GPTQModel=LLaMA2-13B, Precision=W3A16, Evaluation Protocol=Zero-shot2026.05 | 73.88 | — | |
| Cosine SimilarityBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 73.72 | — | |
| BLT-DParameters=1B, Block size=8, Likelihood-based evaluation=true2026.05 | 73.56 | — | |
| BLT-DParameters=1B, Block size=16, Likelihood-based evaluation=true2026.05 | 72.36 | — | |
| PerplexityBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 71.71 | — | |
| PrismScale=Large, Evaluation Protocol=Zero-shot2026.06 | 71.3 | — | |
| BaselineScale=Large, Evaluation Protocol=Zero-shot2026.06 | 70.2 | — | |
| PrismScale=Medium, Evaluation Protocol=Zero-shot2026.06 | 68.2 | — | |
| BaselineScale=Medium, Evaluation Protocol=Zero-shot2026.06 | 67.7 | — | |
| TaylorBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 65.94 | — | |
| PrismScale=Small, Evaluation Protocol=Zero-shot2026.06 | 62.8 | — | |
| BaselineScale=Small, Evaluation Protocol=Zero-shot2026.06 | 62.5 | — | |
| Nautile-370MTraining tokens=∼0.8T, Evaluation Protocol=0-shot2026.04 | 61.5 | — | |
| Qwen2.5 0.5BTraining tokens=18T, Evaluation Protocol=0-shot2026.04 | 61.3 | — | |
| Baselineshot=0-shot, eval-framework=lm-evaluation-harness2026.04 | 59.47 | 59.58 | |
| GMT v7shot=0-shot, eval-framework=lm-evaluation-harness2026.04 | 57.78 | 57.89 | |
| Granite 350MTraining tokens=10–12T, Evaluation Protocol=0-shot2026.04 | 50.8 | — | |
| LFM2.5 350MTraining tokens=28T, Evaluation Protocol=0-shot2026.04 | 49.5 | — | |
| SmolLM2 360MTraining tokens=4T, Evaluation Protocol=0-shot2026.04 | 48.2 | — |