VLM Inference Efficiency on Qwen2.5-VL-7B
643.68Prefill Latency (ms)MBQ
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| MBQQuantization Precision=W4A4, Hardware=RTX 4090, Sequence Length=2048, Batch Size=82026.05 | 643.68 | 3.33 | 8.89 | 1.94 | |
| MASQuantQuantization Precision=W4A4, Hardware=RTX 4090, Sequence Length=2048, Batch Size=82026.05 | 696.44 | 3.07 | 9.42 | 1.83 | |
| SplitQQuantization Precision=W4A4, Hardware=RTX 4090, Sequence Length=2048, Batch Size=82026.05 | 742.17 | 2.89 | 9.48 | 1.82 | |
| FP16Quantization Precision=FP16, Hardware=RTX 4090, Sequence Length=2048, Batch Size=82026.05 | 2,146.01 | — | 17.27 | — |