Image Captioning on NoCaps (NC, Avg %)
43.5NC ScoreLightKV
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LightKVModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 43.5 | 101.37 | |
| PiToMeModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 43.3 | 100.24 | |
| ToMeModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 42.5 | 100.04 | |
| ToFuModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 41.8 | 100.75 | |
| FastVModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 38.6 | 98.77 | |
| VanillaModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 37.2 | 100 |