ResearchBenchmarksInference Latency on LLaMA 2 70BFollow1,450Latency (ms)ROSE1,436.361,528.431,620.51,712.57Mar 6, 2026Evaluation ResultsMethodMethodLinksLatency (ms)SpeedupROSESparsity pattern=2:4,...Sparsity pattern=2:4, Inference kernel=NVIDIA's CUTLASS2026.031,4501.24SparseGPTSparsity pattern=2:4,...Sparsity pattern=2:4, Inference kernel=NVIDIA's CUTLASS2026.031,4581.23DenseSparsity pattern=denseSparsity pattern=dense2026.031,791—