Latency Measurement on Synthetic sequences performance benchmarking
0.03Latency (ms)Topological Attention
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Topological AttentionSeq Length=1,024, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 0.03 | 1.3 | 93.7 | |
| Flash AttentionSeq Length=1,024, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 0.04 | — | — | |
| Topological AttentionSeq Length=8,192, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 0.19 | 16.8 | 99.2 | |
| Topological AttentionSeq Length=65,536, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 2.78 | 56.1 | 99.9 | |
| Flash AttentionSeq Length=8,192, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 3.27 | — | — | |
| Topological AttentionSeq Length=131,072, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 3.94 | 159.3 | 100 | |
| Topological AttentionSeq Length=262,144, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 7.88 | 319 | 100 | |
| Topological AttentionSeq Length=524,288, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 15.76 | 637 | 100 | |
| Topological AttentionSeq Length=1,048,576, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 31.52 | — | 100 | |
| Flash AttentionSeq Length=65,536, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 156.25 | — | — | |
| Flash AttentionSeq Length=131,072, Hardware=NVIDIA RTX 4090 Laptop (16GB VRAM), Batch size=1, Attention config=8 heads, 64 dimensions per head, Window size=64 tokens2026.02 | 627.58 | — | — |