ResearchBenchmarksLLM Inference on LLaMA2-7B 1,024 tokensFollow20Latency (ms)Thor-U14.6450.8287123.18Apr 20, 2026Evaluation ResultsMethodMethodLinksLatency (ms)Speedup FactorThor-UPhase=decode, Quantiza...Phase=decode, Quantization=W4A162026.0420—M100Phase=decode, Quantiza...Phase=decode, Quantization=W4A16, Number of active clusters=122026.0421.340.94M100Phase=prefill, Quantiz...Phase=prefill, Quantization=W8A8, Number of active clusters=122026.04791.95Thor-UPhase=prefill, Quantiz...Phase=prefill, Quantization=W8A82026.04154—