Language Modeling on OpenWebText GPT-2 124M (train)
2.9552LossMVN-GradW
Evaluation Results
| Method | Links | |
|---|---|---|
| MVN-GradWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | 2.9552 | |
| AdaBeliefWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | 2.9558 | |
| AdamWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | 2.9715 | |
| LaPropWBackbone=GPT-2 124M, Block Size=1024, Effective Batch Size=480, Precision=bf16, Weight Decay=Decoupled (0.1)2026.02 | 2.972 |