Language Modeling on WikiText-2 raw-v1 (val)
2.744Cross Entropy (CE)OPT-125M (baseline)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OPT-125M (baseline)Params (M)=125, Notes=Raw WT2 baseline (Table 1, sparsity 0.0). [22]2026.03 | 2.744 | 15.55 | |
| GPT-2 Large (pretrained)Params (M)=774, Notes=WT2-raw-v1; no overlap (stride=1024). 512 stride: 16.44. [21]2026.03 | 2.967 | 19.44 | |
| Pythia–70M (scratch, FP16)Params (M)=70, Notes=low-bit study FP16 reference; CE ≈ 4.29 (protocol differs)2026.03 | 4.298 | 73.1 | |
| SFT Pythia–70M (HH)Params (M)=70, Notes=small SFT model; WT2 word perplexity from model card (split unspecified); CE ≈ 5.192026.03 | 5.195 | 180.27 | |
| Fuzzy-Gated + RPAParams (M)=∼90, Notes=WT2 (raw-v1, GPT-2 BPE); sequential, no-overlap eval2026.03 | 5.246 | 189.8 |