Reasoning on 8-Component Commonsense Suite (ARC-e, ARC-c, BoolQ, PIQA, SIQA, HellaSwag, OBQA, WinoGrande)
72.6ARC-e AccuracyFP16 Baseline
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| FP16 BaselineBackbone Model=LLaMA3-3B, Precision=FP16, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 72.6 | 50.7 | 74.6 | 78.2 | 48.5 | 74.3 | 53.7 | 69.2 | 65.2 | |
| FP16 BaselineBackbone=Qwen-1.7B, Quantization=FP16, Evaluation Protocol=Zero-shot2026.05 | 68.9 | 41 | 78.9 | 71.7 | 45.1 | 59.6 | 38.5 | 61.6 | 58.2 | |
| ParetoQBackbone Model=LLaMA3-3B, Precision=W1.58A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 67.9 | 42 | 53.9 | 72 | 43.9 | 58.2 | 49.6 | 60.2 | 56 | |
| ParetoQ + JacQuantBackbone Model=LLaMA3-3B, Precision=W1.58A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 67.3 | 42 | 57.5 | 73.4 | 43.8 | 59.1 | 51.8 | 61.3 | 57.1 | |
| WinQ + JacQuantBackbone=LLaMA3-1B, Weight bit-width=2, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 66.4 | 39.8 | 64.7 | 73.3 | 43.8 | 59.2 | 47.8 | 59.4 | 56.8 | |
| WinQ + JacQuantBackbone=LLaMA3-1B, Weight bit-width=1.58, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 66.1 | 39.2 | 63.4 | 73.6 | 44.8 | 56.9 | 47.5 | 57.1 | 56.1 | |
| ParetoQ + JacQuantBackbone=LLaMA3-1B, Weight bit-width=2, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 66.1 | 42 | 62.1 | 72.8 | 43.3 | 57.4 | 49.8 | 59.1 | 56.6 | |
| WinQBackbone=LLaMA3-1B, Weight bit-width=1.58, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 65 | 39.7 | 62.1 | 72.9 | 44.3 | 56.2 | 47.1 | 57.4 | 55.6 | |
| WinQBackbone=LLaMA3-1B, Weight bit-width=2, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 65 | 42.5 | 62.5 | 73.8 | 43.2 | 58.4 | 48.1 | 59.1 | 56.6 | |
| FP16 BaselineBackbone=LLaMA3-1B, Weight bit-width=16, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 64.8 | 42.5 | 64.8 | 74.8 | 44.8 | 64.4 | 50.2 | 61.5 | 58.5 | |
| ParetoQ + JacQuantBackbone Model=LLaMA3-3B, Precision=W1A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 64.3 | 38.9 | 60.6 | 72.2 | 43.2 | 50.7 | 49.8 | 59.1 | 54.9 | |
| ParetoQ + JacQuantBackbone=LLaMA3-1B, Weight bit-width=1.58, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 64.2 | 38.2 | 62.2 | 71.6 | 44.4 | 55 | 47.7 | 59.1 | 55.3 | |
| ParetoQBackbone Model=LLaMA3-3B, Precision=W1A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 63.1 | 39.1 | 63 | 70.9 | 42.6 | 50.6 | 46.7 | 56.5 | 54.1 | |
| ParetoQBackbone=LLaMA3-1B, Weight bit-width=2, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 63.1 | 41.9 | 62.5 | 72.8 | 43.6 | 55.2 | 49.2 | 57.3 | 55.7 | |
| ParetoQBackbone=LLaMA3-1B, Weight bit-width=1.58, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 61.4 | 39.7 | 63.1 | 71 | 42.6 | 53.1 | 47.5 | 55.7 | 54.3 | |
| WinQ + JacQuantBackbone=LLaMA3-1B, Weight bit-width=1, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 61.1 | 37.9 | 63.1 | 70.2 | 43.2 | 50.9 | 44.9 | 56.3 | 53.5 | |
| ParetoQ + JacQuantBackbone=LLaMA3-1B, Weight bit-width=1, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 60.6 | 38.9 | 64.1 | 67.2 | 42.6 | 47.8 | 44.1 | 53.9 | 52.4 | |
| WinQBackbone=LLaMA3-1B, Weight bit-width=1, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 59.7 | 37 | 61.3 | 69.5 | 42.6 | 49.8 | 46.1 | 54.4 | 52.6 | |
| ParetoQBackbone=LLaMA3-1B, Weight bit-width=1, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 58.7 | 39.8 | 62.6 | 67.4 | 42.4 | 46.9 | 43.9 | 53.5 | 51.9 | |
| JacQuantBackbone=Qwen-1.7B, Quantization=W2A8, Evaluation Protocol=Zero-shot, Training Dataset=FineWebEdu, Training Method=QAT, Baseline=ParetoQ2026.05 | 54.1 | 31.1 | 64.4 | 64.1 | 43.5 | 40.3 | 34.1 | 52.8 | 48.1 | |
| ParetoQBackbone=Qwen-1.7B, Quantization=W2A8, Evaluation Protocol=Zero-shot, Training Dataset=FineWebEdu, Training Method=QAT2026.05 | 52.9 | 32.6 | 62.4 | 65.1 | 42.7 | 41.2 | 32.8 | 53.1 | 47.9 | |
| JacQuantBackbone=Qwen-1.7B, Quantization=W1A8, Evaluation Protocol=Zero-shot, Training Dataset=FineWebEdu, Training Method=QAT, Baseline=ParetoQ2026.05 | 42.5 | 26.5 | 61.8 | 59.6 | 40.6 | 29.5 | 29.4 | 51.7 | 42.7 | |
| ParetoQBackbone=Qwen-1.7B, Quantization=W1A8, Evaluation Protocol=Zero-shot, Training Dataset=FineWebEdu, Training Method=QAT2026.05 | 42 | 26.2 | 61.4 | 59.8 | 40 | 29.1 | 28.9 | 50.8 | 42.3 | |
| SpinQuantBackbone=LLaMA3-1B, Weight bit-width=2, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 31 | 20.1 | 45.5 | 54.6 | 33.8 | 26.7 | 16.2 | 50.9 | 34.9 | |
| SpinQuantBackbone Model=LLaMA3-3B, Precision=W1A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 27 | 22 | 37.6 | 52 | 33.4 | 25.5 | 14.5 | 50.2 | 32.8 | |
| GPTQBackbone=LLaMA3-1B, Weight bit-width=1, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 26.9 | 21.7 | 37.6 | 51.8 | 33.5 | 25.5 | 14.8 | 49.7 | 32.7 | |
| GPTQBackbone=LLaMA3-1B, Weight bit-width=2, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 26.9 | 21.7 | 37.6 | 51.8 | 33.5 | 25.5 | 14.8 | 49.7 | 32.7 | |
| GPTQBackbone Model=LLaMA3-3B, Precision=W1A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 26.6 | 21.9 | 37.6 | 52.5 | 33.7 | 25.6 | 15 | 49.4 | 32.8 | |
| SpinQuantBackbone Model=LLaMA3-3B, Precision=W1.58A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 26.6 | 23.4 | 39.1 | 52.7 | 34.7 | 25.3 | 16.4 | 48.5 | 33.3 | |
| GPTQBackbone Model=LLaMA3-3B, Precision=W1.58A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 25.3 | 22.9 | 37.6 | 53.7 | 34.4 | 25.3 | 15.6 | 49.5 | 33 | |
| SpinQuantBackbone=LLaMA3-1B, Weight bit-width=1.58, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 25.3 | 22.5 | 37.6 | 51.6 | 33.3 | 25.3 | 17.6 | 48.5 | 32.7 | |
| RTNBackbone Model=LLaMA3-3B, Precision=W1A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 25 | 22.5 | 37.6 | 49.5 | 32.9 | 25 | 27.1 | 49.6 | 33.7 | |
| RTNBackbone=LLaMA3-1B, Weight bit-width=1, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 25 | 22.5 | 37.6 | 49.5 | 32.9 | 25 | 27.1 | 49.6 | 33.7 | |
| SpinQuantBackbone=LLaMA3-1B, Weight bit-width=1, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 25 | 22.5 | 37.6 | 49.5 | 32.9 | 25 | 27.1 | 49.6 | 33.7 | |
| GPTQBackbone=LLaMA3-1B, Weight bit-width=1.58, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 24.8 | 22.2 | 38 | 52.2 | 32.5 | 25.4 | 17 | 49.4 | 32.7 | |
| RTNBackbone Model=LLaMA3-3B, Precision=W1.58A8, Evaluation Protocol=Zero-shot, Training Context=QAT on FineWebEdu2026.05 | 24.7 | 23.4 | 37.6 | 52.9 | 33.8 | 25.5 | 17.4 | 50.3 | 33.2 | |
| RTNBackbone=LLaMA3-1B, Weight bit-width=1.58, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 24.5 | 22.3 | 62.4 | 52.7 | 33.4 | 25.4 | 18.4 | 50.2 | 36.2 | |
| RTNBackbone=LLaMA3-1B, Weight bit-width=2, Activation bit-width=16, Evaluation protocol=Zero-shot2026.05 | 24.5 | 23.1 | 62.4 | 52.3 | 33.6 | 25.4 | 17.6 | 50.3 | 36.2 |