Image Captioning on MS COCO (Coco, Avg %)
38.9COCO ScorePiToMe
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PiToMeModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 38.9 | 100.24 | |
| LightKVModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 38.9 | 101.37 | |
| ToFuModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 38.3 | 100.75 | |
| FastVModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 33.9 | 98.77 | |
| ToMeModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 32.9 | 100.04 | |
| VanillaModel=Qwen2.5-VL-7B-Instruct, Vision token retention rate=55%2026.05 | 31.9 | 100 |