Commonsense Reasoning on HellaSwag (HS, Avg.)
78.9HS AccuracyFP16
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FP16Bits=16.02026.07 | 78.9 | — | |
| TASA b4.0Bits=4.02026.07 | 78.6 | — | |
| SpQR†Bits=4.12 (W3)2026.07 | 78.3 | — | |
| AWQBits=4.02026.07 | 78.2 | — | |
| HQQBits=4.02026.07 | 78.2 | — | |
| GPTQBits=4.02026.07 | 78.1 | — | |
| RTNBits=4.02026.07 | 78 | — | |
| DenseModel=LLaMA-2-7B, Sparsity=0%, Evaluation=Zero-shot2026.07 | 76.5 | — | |
| Qwen3Model Version=4B, Architecture=Dense, # Params=4.0B, # Tokens=36T2025.10 | 75.66 | — | |
| Gemma3Model Version=4B, Architecture=Dense, # Params=4.0B, # Tokens=4T2025.10 | 75.58 | — | |
| OWQBits=3.02026.07 | 75.4 | — | |
| TASA b3.0Bits=3.02026.07 | 75.4 | — | |
| Qwen-4B (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 75 | 66.42 | |
| HQQBits=3.02026.07 | 75 | — | |
| Llama-3B (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 74.92 | 46.93 | |
| Qwen2.5Model Version=3B, Architecture=Dense, # Params=3.0B, # Tokens=18T2025.10 | 74.54 | — | |
| OuroModel Version=1.4B R4, Architecture=LoopLM, # Params=1.4B, # Tokens=7.7T, recurrent steps=42025.10 | 74.29 | — | |
| GPTQBits=3.02026.07 | 73.9 | — | |
| Phi-mini (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 73.36 | 63.72 | |
| Llama3.2Model Version=3B, Architecture=Dense, # Params=3.0B, # Tokens=9T2025.10 | 73.09 | — | |
| PALSModel=LLaMA-2-7B, Sparsity=50%, Evaluation=Zero-shot2026.07 | 71.4 | — | |
| WandaModel=LLaMA-2-7B, Sparsity=50%, Evaluation=Zero-shot2026.07 | 70.2 | — | |
| Qwen2.5Model Version=1.5B, Architecture=Dense, # Params=1.5B, # Tokens=18T2025.10 | 67.73 | — | |
| Qwen3Model Version=1.7B, Architecture=Dense, # Params=1.7B, # Tokens=36T2025.10 | 67.09 | — | |
| Llama-1B (base)Setting=No distillation, Few-shot protocol=3-shot2026.05 | 65.08 | 33.96 | |
| Llama-3B → 1BSetting=Same tokenizer, Teacher=Llama-3B, Few-shot protocol=3-shot2026.05 | 64.42 | 38.4 | |
| Continued pre-trainingSetting=No distillation, Few-shot protocol=3-shot2026.05 | 63.9 | 36.63 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Qwen-4B + Llama-3B, Few-shot protocol=3-shot2026.05 | 63.55 | 40.15 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Llama-3B, Few-shot protocol=3-shot2026.05 | 63.38 | 40.48 | |
| ULDSetting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 63.32 | 38.31 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Qwen-4B, Few-shot protocol=3-shot2026.05 | 63 | 38.49 | |
| ULDSetting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 62.93 | 36.77 | |
| GOLDSetting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 62.92 | 38.66 | |
| X-Token (H-KL)Setting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 62.67 | 39.18 | |
| X-Token (P-KL)Setting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 62.63 | 38.85 | |
| GOLDSetting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 62.59 | 35.03 | |
| MagnitudeModel=LLaMA-2-7B, Sparsity=50%, Evaluation=Zero-shot2026.07 | 61.6 | — | |
| Llama3.2Model Version=1.2B, Architecture=Dense, # Params=1.0B, # Tokens=9T2025.10 | 59.35 | — | |
| Gemma3Model Version=1B, Architecture=Dense, # Params=1.0B, # Tokens=2T2025.10 | 56.12 | — | |
| RTNBits=3.02026.07 | 49.6 | — | |
| ADAMS (DENSE all-reduce)Protocol=Zero-shot, Model=GPT-345M, Strategy=DENSE all-reduce2026.07 | 31.45 | — | |
| SCAPE (d = 0.1)Protocol=Zero-shot, Model=GPT-345M, d=0.12026.07 | 31.28 | — | |
| SCAPE (d = 0.01)Protocol=Zero-shot, Model=GPT-345M, d=0.012026.07 | 31.06 | — | |
| ADAMW (DENSE all-reduce)Protocol=Zero-shot, Model=GPT-345M, Strategy=DENSE all-reduce2026.07 | 30.63 | — | |
| Gemma3Architecture=Dense, # Total Params=4.0B, # Trained Tokens=4T2025.10 | — | 75.58 | |
| Gemma3Architecture=Dense, # Total Params=12.0B, # Trained Tokens=12T2025.10 | — | 83.68 | |
| Llama3.1Architecture=Dense, # Total Params=8.0B, # Trained Tokens=15T2025.10 | — | 81.97 | |
| Llama3.2Architecture=Dense, # Total Params=3.0B, # Trained Tokens=9T2025.10 | — | 73.09 | |
| Ouro 2.6B R4Architecture=LoopLM, # Total Params=2.6B, # Trained Tokens=7.7T2025.10 | — | 79.69 | |
| Qwen2.5Architecture=Dense, # Total Params=3.0B, # Trained Tokens=18T2025.10 | — | 74.54 | |
| Qwen2.5Architecture=Dense, # Total Params=7.0B, # Trained Tokens=18T2025.10 | — | 79.98 | |
| Qwen3Architecture=Dense, # Total Params=4.0B, # Trained Tokens=36T2025.10 | — | 75.66 | |
| Qwen3Architecture=Dense, # Total Params=8.0B, # Trained Tokens=36T2025.10 | — | 79.6 |