Decoding on RTX 6000 Ada platform
468.7Throughput (tokens/sec)LAQuant
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LAQuantModel=Qwen3-1.7B, Bits=3, Batch size=12026.05 | 468.7 | 2.5 | |
| ParoQuantModel=Qwen3-1.7B, Bits=3, Batch size=12026.05 | 392.9 | 2.1 | |
| LAQuantModel=Llama3.1-8B, Bits=3, Batch size=12026.05 | 345.47 | 6.45 | |
| LAQuantModel=Qwen3-4B, Bits=3, Batch size=12026.05 | 325 | 3.58 | |
| LAQuantModel=Qwen3-1.7B, Bits=4, Batch size=12026.05 | 313.9 | 1.68 | |
| ParoQuantModel=Llama3.1-8B, Bits=3, Batch size=12026.05 | 298.15 | 5.57 | |
| ParoQuantModel=Qwen3-1.7B, Bits=4, Batch size=12026.05 | 279.7 | 1.49 | |
| ParoQuantModel=Qwen3-4B, Bits=3, Batch size=12026.05 | 278.2 | 3.06 | |
| LAQuantModel=Qwen3-8B, Bits=3, Batch size=12026.05 | 276.9 | 5.27 | |
| ParoQuantModel=Qwen3-8B, Bits=3, Batch size=12026.05 | 241.3 | 4.59 | |
| LAQuantModel=Qwen3-4B, Bits=4, Batch size=12026.05 | 191.4 | 2.11 | |
| FP16Model=Qwen3-1.7B, Bits=16, Batch size=12026.05 | 187.1 | — | |
| ParoQuantModel=Qwen3-4B, Bits=4, Batch size=12026.05 | 174.4 | 1.92 | |
| LAQuantModel=Llama3.1-8B, Bits=4, Batch size=12026.05 | 129.26 | 2.41 | |
| LAQuantModel=Qwen3-8B, Bits=4, Batch size=12026.05 | 122.5 | 2.33 | |
| ParoQuantModel=Llama3.1-8B, Bits=4, Batch size=12026.05 | 120.93 | 2.26 | |
| ParoQuantModel=Qwen3-8B, Bits=4, Batch size=12026.05 | 112.6 | 2.14 | |
| FP16Model=Qwen3-4B, Bits=16, Batch size=12026.05 | 90.8 | — | |
| FP16Model=Llama3.1-8B, Bits=16, Batch size=12026.05 | 53.55 | — | |
| FP16Model=Qwen3-8B, Bits=16, Batch size=12026.05 | 52.5 | — |