Language Modeling on Lambada S
18.43AccuracyADAPT-BM25
Evaluation Results
| Method | Links | |
|---|---|---|
| ADAPT-BM25Backbone=TinyLlama-120M, Training Budget=50B tokens, FLOPs overhead=≪ 1.0 × 10^142026.04 | 18.43 | |
| RegMixBackbone=TinyLlama-120M, Training Budget=50B tokens, FLOPs overhead=3.072 × 10^182026.04 | 18.3 | |
| ADAPTBackbone=TinyLlama-120M, Training Budget=50B tokens, FLOPs overhead=≪ 1.1 × 10^152026.04 | 18.07 | |
| UniformBackbone=TinyLlama-120M, Training Budget=50B tokens, FLOPs overhead=02026.04 | 16.98 | |
| LinUpperBackbone=TinyLlama-120M, Training Budget=50B tokens, FLOPs overhead=02026.04 | 16.79 | |
| DoReMiBackbone=TinyLlama-120M, Training Budget=50B tokens, FLOPs overhead=4.92 × 10^192026.04 | 16.3 |