Decoding Performance on RTX 4090 platform
426Throughput (tokens/sec)LAQuant
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LAQuantModel=Qwen3-1.7B, Bits=3, Batch size=12026.05 | 426 | 2.27 | |
| ParoQuantModel=Qwen3-1.7B, Bits=3, Batch size=12026.05 | 365.8 | 1.95 | |
| LAQuantModel=Llama3.1-8B, Bits=3, Batch size=12026.05 | 322.11 | 5.72 | |
| LAQuantModel=Qwen3-4B, Bits=3, Batch size=12026.05 | 322.1 | 3.48 | |
| LAQuantModel=Qwen3-1.7B, Bits=4, Batch size=12026.05 | 302.1 | 1.61 | |
| ParoQuantModel=Llama3.1-8B, Bits=3, Batch size=12026.05 | 282.94 | 5.3 | |
| ParoQuantModel=Qwen3-4B, Bits=3, Batch size=12026.05 | 279.6 | 3.02 | |
| LAQuantModel=Qwen3-8B, Bits=3, Batch size=12026.05 | 276.6 | 5.03 | |
| ParoQuantModel=Qwen3-1.7B, Bits=4, Batch size=12026.05 | 271.6 | 1.45 | |
| ParoQuantModel=Qwen3-8B, Bits=3, Batch size=12026.05 | 243.2 | 4.42 | |
| FP16Model=Qwen3-1.7B, Bits=16, Batch size=12026.05 | 187.3 | — | |
| LAQuantModel=Qwen3-4B, Bits=4, Batch size=12026.05 | 183.5 | 1.99 | |
| ParoQuantModel=Qwen3-4B, Bits=4, Batch size=12026.05 | 168.4 | 1.82 | |
| LAQuantModel=Llama3.1-8B, Bits=4, Batch size=12026.05 | 129.71 | 2.3 | |
| ParoQuantModel=Llama3.1-8B, Bits=4, Batch size=12026.05 | 122.66 | 2.18 | |
| LAQuantModel=Qwen3-8B, Bits=4, Batch size=12026.05 | 121.8 | 2.21 | |
| ParoQuantModel=Qwen3-8B, Bits=4, Batch size=12026.05 | 114.6 | 2.08 | |
| FP16Model=Qwen3-4B, Bits=16, Batch size=12026.05 | 92.5 | — | |
| FP16Model=Llama3.1-8B, Bits=16, Batch size=12026.05 | 56.3 | — | |
| FP16Model=Qwen3-8B, Bits=16, Batch size=12026.05 | 55 | — |