Multimodal Large Language Model Inference on Qwen2VL-7B FP16 (inference)
7.13Power Consumption (W)SpikeMLLM (Proposed accelerator)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SpikeMLLM (Proposed accelerator)System=SpikeMLLM (Proposed accelerator), Batch size=1, Configuration=QuaRot+MSTS+TC-LIF, Technology=SMIC 28nm2026.04 | 7.13 | 393.7 | |
| Qwen2VL-7B-FP16 (NVIDIA A800)System=NVIDIA A800, Batch size=1, Precision=FP162026.04 | 184 | 43.5 |