Large Language Model Inference on Qwen VL 7B 2.5
9Throughput (Intel i7-13620H)LUQ
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LUQModel Configuration=LUQ (Mixed Precision), Inference Engine=llama.cpp, Precision=Mixed Precision2025.09 | 9 | 18.7 | 3.4 | |
| Q4_K_MModel Configuration=Q4_K_M (Standard 4-bit), Inference Engine=llama.cpp, Quantization Level=4-bit2025.09 | 4.8 | 14.1 | 4.4 | |
| FP16Model Configuration=FP16 (Baseline), Inference Engine=llama.cpp, Precision=FP162025.09 | 0.2 | 4.9 | 14.5 |