Language Modeling on OWT Pile
5.23Decode LatencyLongformer + SFA (k=8)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Longformer + SFA (k=8)Backbone=GPT-2 124M, Method Category=Token Sparse (Training), Latency Context Length=128k, Evaluation Protocol=Training, SFA k parameter=82026.03 | 5.23 | 6.18 | 19.3 | |
| LongformerBackbone=GPT-2 124M, Method Category=Token Sparse (Training), Latency Context Length=128k, Evaluation Protocol=Training2026.03 | 6.75 | 7.93 | 18.73 | |
| SnapKV + SFA (k=8)Backbone=GPT-2 124M, Method Category=KV-pruning (Training-free), Latency Context Length=128k, Evaluation Protocol=Training-free, SFA k parameter=82026.03 | 6.92 | 9.41 | 19.44 | |
| RoutingBackbone=GPT-2 124M, Method Category=Token Sparse (Training), Latency Context Length=128k, Evaluation Protocol=Training2026.03 | 7.92 | 8.37 | 18.64 | |
| Loki + SFA (k=8)Backbone=GPT-2 124M, Method Category=Low-rank keys (Training-free), Latency Context Length=128k, Evaluation Protocol=Training-free, SFA k parameter=82026.03 | 9.09 | 9.41 | 19.29 | |
| PerformerBackbone=GPT-2 124M, Method Category=Kernel Method, Latency Context Length=128k2026.03 | 9.43 | 7.93 | 19.72 | |
| SnapKVBackbone=GPT-2 124M, Method Category=KV-pruning (Training-free), Latency Context Length=128k, Evaluation Protocol=Training-free2026.03 | 9.88 | 16.86 | 17.91 | |
| QuestBackbone=GPT-2 124M, Method Category=KV-pruning (Training-free), Latency Context Length=128k, Evaluation Protocol=Training-free2026.03 | 10.84 | 16.86 | 17.95 | |
| LokiBackbone=GPT-2 124M, Method Category=Low-rank keys (Training-free), Latency Context Length=128k, Evaluation Protocol=Training-free2026.03 | 11.39 | 16.86 | 17.82 | |
| H2OBackbone=GPT-2 124M, Method Category=KV-pruning (Training-free), Latency Context Length=128k, Evaluation Protocol=Training-free2026.03 | 13.32 | 16.86 | 18.02 | |
| SFABackbone=GPT-2 124M, Method Category=Feature Sparse Attention, Latency Context Length=128k2026.03 | 14.12 | 9.41 | 18.17 | |
| Dense (full)Backbone=GPT-2 124M, Method Category=Full Attention Baseline, Latency Context Length=128k2026.03 | 17.08 | 16.86 | 17.29 |