Language Modeling on OpenWebText GPT-2 124M (val)
3.167LCEBaseline float32
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Baseline float32N bits=32 bits, quantization_protocol=PTQ2026.01 | 3.167 | 23.73 | — | |
| Baseline float32N bits=32 bits, quantization_protocol=QAT2026.01 | 3.167 | 23.167 | — | |
| PoT - QATInteger range=[-2^7, 2^7], N bits=4 bits2026.01 | 3.417 | 30.47 | — | |
| PoT - PTQInteger range=[-2^7, 2^7], N bits=4 bits2026.01 | 4.5 | 90 | — | |
| PoT - QATInteger range=[-2^5, 2^5], N bits=4 bits2026.01 | 4.71 | 111 | — | |
| PoT - QATInteger range=[-2^3, 2^3], N bits=3 bits2026.01 | 7.42 | 1,669 | — | |
| PoT - PTQInteger range=[-2^5, 2^5], N bits=4 bits2026.01 | 8.76 | 6,374 | — | |
| PoT - PTQInteger range=[-2^3, 2^3], N bits=3 bits2026.01 | 9.01 | 8,184 | — | |
| AdaBeliefWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | — | — | 2.9744 | |
| AdamWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | — | — | 2.9897 | |
| LaPropWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | — | — | 2.9904 | |
| MVN-GradWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | — | — | 2.9735 |