Inference Efficiency on 1x V100 (16GB) (synthetic)
28,298Throughput (tokens/s)SRM
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SRMn_ctx=1024, d_m=1024, n_l=16, h=4, Implementation=Pytorch2026.05 | 28,298 | 64,000 | 14.75 | |
| SRMn_ctx=512, d_m=1024, n_l=16, h=4, Implementation=Pytorch2026.05 | 28,091 | 64,000 | 9.66 | |
| SRMn_ctx=4096, d_m=1024, n_l=16, h=4, Implementation=Pytorch2026.05 | 27,445 | 32,000 | 43.29 | |
| SRMn_ctx=2048, d_m=1024, n_l=16, h=4, Implementation=Pytorch2026.05 | 27,441 | 32,000 | 24.59 | |
| Transformern_ctx=512, d_m=512, n_l=16, h=4, Implementation=Pytorch2026.05 | 2,908 | 400 | — | |
| Transformern_ctx=1024, d_m=512, n_l=16, h=4, Implementation=Pytorch2026.05 | 1,918 | 200 | — | |
| Transformern_ctx=2048, d_m=512, n_l=16, h=4, Implementation=Pytorch2026.05 | 1,116 | 100 | — | |
| Transformern_ctx=4096, d_m=512, n_l=16, h=4, Implementation=Pytorch2026.05 | 634 | 50 | — |