Zero-shot Language Modeling and Reasoning on PIQA, ARC, HellaSwag, WinoG, BoolQ, LAMBADA, and C4
76.33PIQA AccuracyL2QER
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| L2QERRank=64, Evaluation=Zero-shot2026.03 | 76.33 | 42.33 | 70 | 66 | 67 | 80.67 | 68 | 8.93 | 67.19 | |
| FP16Rank=-, Evaluation=Zero-shot2026.03 | 76 | 41.33 | 70.33 | 66 | 68 | 80.33 | 72.33 | 8.7 | 67.76 | |
| QERARank=64, Evaluation=Zero-shot2026.03 | 76 | 42.67 | 70.67 | 67 | 67.33 | 81 | 69.67 | 8.91 | 67.76 | |
| GlowQ-SRank=64, Evaluation=Zero-shot2026.03 | 76 | 43.67 | 69.33 | 66 | 66.67 | 82 | 70 | 8.99 | 67.67 | |
| ZeroQuant-V2Rank=64, Evaluation=Zero-shot2026.03 | 75.67 | 43 | 71 | 65 | 65.33 | 80 | 68.33 | 9.07 | 66.9 | |
| GlowQRank=64, Evaluation=Zero-shot2026.03 | 75.67 | 41.67 | 70 | 66.67 | 67 | 80.33 | 69.67 | 8.87 | 67.29 |