Video Captioning on VDC (Component-wise Acc/Sim Metrics)
49.2Short AccuracyVideoZoomer
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VideoZoomer2025.12 | 49.2 | — | — | 2.51 | — | — | — | — | — | — | — | — | |
| Our LFS + Qwen3-VL-8BSampling Strategy=Our LFS, Backbone Model=Qwen3-VL-8B2026.01 | 39.58 | 48.82 | 2.97 | 1.75 | 41.62 | 2.59 | 43.59 | 2.87 | 57.58 | 2.71 | — | — | |
| ShareGPT4Video-8BSampling Strategy=Uniform Sampling, Backbone Model=ShareGPT4Video-8B2026.01 | 39.08 | 33.28 | 1.76 | 1.94 | 35.77 | 1.81 | 37.12 | 1.89 | 35.62 | 1.84 | — | — | |
| Qwen3-VL-8BSampling Strategy=Uniform Sampling, Backbone Model=Qwen3-VL-8B2026.01 | 38.28 | 47.66 | 2.69 | 1.35 | 39.98 | 2.16 | 42.36 | 2.56 | 55.56 | 2.59 | — | — | |
| QwenVL-2.5-7BParameters=7B2025.12 | 37.8 | — | — | 1.98 | — | — | — | — | — | — | — | — | |
| Gemini-1.5 ProSampling Strategy=Not Specified, Backbone Model=Gemini-1.5 Pro2026.01 | 35.71 | 38.68 | 2.05 | 1.85 | 43.84 | 2.23 | 47.32 | 2.41 | 43.11 | 2.22 | — | — | |
| Our LFS + AuroraCap-7BSampling Strategy=Our LFS, Backbone Model=AuroraCap-7B2026.01 | 34.57 | 44.1 | 2.35 | 1.77 | 36.02 | 1.98 | 40.65 | 2.77 | 43.04 | 2.21 | — | — | |
| InternVL-2-8BSampling Strategy=Uniform Sampling, Backbone Model=InternVL-2-8B2026.01 | 33.02 | 39.08 | 2.11 | 1.74 | 37.47 | 1.89 | 44.16 | 2.22 | 34.89 | 1.82 | — | — | |
| MovieChat-7BSampling Strategy=Uniform Sampling, Backbone Model=MovieChat-7B2026.01 | 32.55 | 37.25 | 1.98 | 1.59 | 28.99 | 1.54 | 31.97 | 1.64 | 28.82 | 1.46 | — | — | |
| AuroraCap-7BSampling Strategy=Uniform Sampling, Backbone Model=AuroraCap-7B2026.01 | 32.07 | 43.5 | 2.27 | 1.68 | 35.92 | 1.84 | 39.02 | 1.97 | 41.3 | 2.15 | — | — | |
| LongVA-7BSampling Strategy=Uniform Sampling, Backbone Model=LongVA-7B2026.01 | 31.94 | 35.32 | 1.9 | 1.63 | 36.39 | 1.85 | 40.95 | 2.11 | 27.91 | 1.48 | — | — | |
| Video-LLaVA-7BSampling Strategy=Uniform Sampling, Backbone Model=Video-LLaVA-7B2026.01 | 30.67 | 37.48 | 1.97 | 1.63 | 32.5 | 1.7 | 36.01 | 1.85 | 27.36 | 1.43 | — | — | |
| VILA-7BSampling Strategy=Uniform Sampling, Backbone Model=VILA-7B2026.01 | 30.4 | 34.33 | 1.83 | 1.55 | 35.15 | 1.8 | 33.38 | 1.72 | 29.78 | 1.58 | — | — | |
| LLaMA-VIDSampling Strategy=Uniform Sampling, Backbone Model=LLaMA-VID2026.01 | 29.92 | 39.47 | 2.1 | 1.56 | 28.01 | 1.45 | 31.24 | 1.59 | 25.67 | 1.38 | — | — | |
| Video-ChatGPT-7BSampling Strategy=Uniform Sampling, Backbone Model=Video-ChatGPT-7B2026.01 | 29.36 | 37.46 | 2 | 1.56 | 33.68 | 1.7 | 30.47 | 1.6 | 24.61 | 1.26 | — | — | |
| ASID-CaptionerSize=7B2026.02 | 28.8 | 38.2 | 1.7 | 1.3 | 46.9 | 2.1 | 47.4 | 2.1 | 43.2 | 1.9 | 40.9 | 1.8 | |
| LLaVA-1.5-7BSampling Strategy=Uniform Sampling, Backbone Model=LLaVA-1.5-7B2026.01 | 28.61 | 38.38 | 2.04 | 1.51 | 34.86 | 1.79 | 34.62 | 1.76 | 33.43 | 1.73 | — | — | |
| ASID-CaptionerSize=3B2026.02 | 28.4 | 37 | 1.6 | 1.3 | 45.5 | 2 | 44.7 | 2 | 41.7 | 1.8 | 39.5 | 1.7 | |
| Gemini-3-ProSize=-2026.02 | 27.6 | 35.5 | 1.5 | 1.2 | 42.3 | 1.9 | 45.4 | 2 | 40.2 | 1.8 | 38.2 | 1.7 | |
| AvoCaDOSize=7B2026.02 | 27.1 | 35.3 | 1.5 | 1.2 | 41 | 1.8 | 40 | 1.8 | 37.9 | 1.7 | 36.3 | 1.6 | |
| Qwen3-Omni-CaptionerSize=30B-A3B2026.02 | 26.8 | 34.4 | 1.4 | 1.2 | 42.1 | 1.9 | 44.7 | 2 | 39.4 | 1.7 | 37.5 | 1.6 | |
| Gemini-2.5-ProSize=-2026.02 | 26.6 | 35 | 1.5 | 1.1 | 41.5 | 1.8 | 42.3 | 1.8 | 38.5 | 1.7 | 36.8 | 1.6 | |
| Gemini-2.5-FlashSize=-2026.02 | 26.5 | 34.5 | 1.4 | 1.1 | 41.1 | 1.7 | 41.6 | 1.8 | 38.1 | 1.7 | 36.4 | 1.5 | |
| Qwen3-Omni-InstructSize=30B-A3B2026.02 | 26.3 | 33.7 | 1.4 | 1.1 | 41.8 | 1.9 | 42.7 | 1.9 | 38.2 | 1.6 | 36.5 | 1.6 | |
| Qwen2.5-OmniSize=7B2026.02 | 26.1 | 29.5 | 1.3 | 1.2 | 38.2 | 1.7 | 37.3 | 1.7 | 34.2 | 1.5 | 33.1 | 1.5 | |
| ARC-Hunyuan-VideoSize=7B2026.02 | 25.5 | 29.1 | 1.3 | 1.2 | 37.4 | 1.6 | 36.5 | 1.5 | 34.3 | 1.5 | 32.6 | 1.4 | |
| Qwen3-VL-235B-Instr2026.05 | 23.56 | 32.87 | — | — | 39.64 | — | 39 | — | 38.57 | — | 34.73 | — | |
| VCapstage=e22026.05 | 23.5 | 34 | — | — | 41.56 | — | 40.74 | — | 40.22 | — | 36.01 | — | |
| VCapstage=e12026.05 | 23.42 | 33.22 | — | — | 40.71 | — | 39.82 | — | 38.9 | — | 35.21 | — | |
| Qwen2.5-OmniSize=3B2026.02 | 23.2 | 28 | 1.3 | 1.2 | 35.1 | 1.6 | 34.6 | 1.7 | 32.4 | 1.4 | 30.7 | 1.4 | |
| Vicuna-v1.5-7BSampling Strategy=Uniform Sampling, Backbone Model=Vicuna-v1.5-7B2026.01 | 23.06 | 21.68 | 1.12 | 1.17 | 22.02 | 1.15 | 22.64 | 1.16 | 23.09 | 1.2 | — | — | |
| Qwen3-VL-8B-Instr2026.05 | 22.96 | 32.54 | — | — | 25.63 | — | 35.28 | — | 37.51 | — | 30.78 | — | |
| Qwen3.5-397B2026.05 | 22.96 | 30.53 | — | — | 37.76 | — | 36.88 | — | 35.87 | — | 32.8 | — | |
| Gemini 3.1 Pro2026.05 | 22.57 | 28.54 | — | — | 36.56 | — | 35.43 | — | 33.98 | — | 31.41 | — | |
| Seed 2.0 Pro2026.05 | 21.54 | 27.39 | — | — | 34.91 | — | 34.55 | — | 33.54 | — | 30.39 | — |