Language Modeling on TinyStories (val)
1.1284Last LossTMMFormer
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TMMFormer2026.05 | 1.1284 | — | — | |
| SOAPFormer2026.05 | 1.1431 | — | — | |
| AdamWFormer2026.05 | 1.1472 | — | — | |
| MuonFormer2026.05 | 1.1503 | — | — | |
| AdamFormer2026.05 | 1.1528 | — | — | |
| VanillaTransformer2026.05 | 1.1569 | — | — | |
| Presymp ETD-AB2wall clock time (seconds)=5798.2, optimization steps=100002026.03 | 1.8386 | 1.8386 | — | |
| Plain Eulerwall clock time (seconds)=5036.1, optimization steps=100002026.03 | 2.3247 | 2.3234 | — | |
| Presymp ExpEulerwall clock time (seconds)=5623.4, optimization steps=100002026.03 | 2.3579 | 2.2523 | — | |
| YuriiFormer-Lie-Trotterwall clock time (seconds)=2835.2, optimization steps=100002026.03 | 2.4041 | 2.3872 | — | |
| Baselinewall clock time (seconds)=2409.4, optimization steps=100002026.03 | 2.4687 | 2.4473 | — | |
| Presymp Eulerwall clock time (seconds)=4979.0, optimization steps=100002026.03 | 2.4728 | 2.4592 | — | |
| Presymp AB2wall clock time (seconds)=5249.4, optimization steps=100002026.03 | 2.6653 | 2.6546 | — | |
| plain Eulerattention=linear, configuration=larger, optimization steps=10000, wall clock time (seconds)=10806.52026.03 | 2.7483 | 2.6567 | — | |
| Presymp Eulerattention=linear, configuration=larger, optimization steps=10000, wall clock time (seconds)=12620.02026.03 | 2.7862 | 2.6909 | — | |
| Presymp Euleroptimization steps=10000, attention mechanism=linear attention, model configuration size=smaller configuration2026.03 | 2.8467 | 2.7314 | 1,172.8 | |
| YuriiFormerattention=linear, configuration=larger, optimization steps=10000, wall clock time (seconds)=9601.32026.03 | 2.8486 | 2.765 | — | |
| YuriiFormeroptimization steps=10000, attention mechanism=linear attention, model configuration size=smaller configuration2026.03 | 2.8783 | 2.766 | 749.6 | |
| Baselineattention=linear, configuration=larger, optimization steps=10000, wall clock time (seconds)=8956.82026.03 | 3.0118 | 2.9172 | — | |
| plain Euleroptimization steps=10000, attention mechanism=linear attention, model configuration size=smaller configuration2026.03 | 3.0399 | 2.9206 | 809.3 | |
| Baselineoptimization steps=10000, attention mechanism=linear attention, model configuration size=smaller configuration2026.03 | 3.083 | 2.9614 | 486.7 |