Video-to-command generation on Modified Something-Something Standard Object Sets V2 (test)
0.626BLEU-1Intern3.5VL-4B
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Intern3.5VL-4B2026.06 | 0.626 | 0.427 | 0.339 | 0.251 | 0.614 | 0.521 | 1.422 | |
| Decoupled Object-Centric Video Understanding framework (Ours Full)Backbone=ResNet-50, Temporal Module=TSM, Detector=HOI detector [11], VLM=LLaVA [8]2026.06 | 0.607 | 0.506 | 0.397 | 0.337 | 0.598 | 0.544 | 1.812 | |
| NVILA-8B2026.06 | 0.557 | 0.389 | 0.329 | 0.286 | 0.561 | 0.475 | 1.773 | |
| Qwen2.5-VL-3B2026.06 | 0.531 | 0.322 | 0.254 | 0.207 | 0.547 | 0.428 | 1.307 | |
| Intern2.5VL-4B2026.06 | 0.525 | 0.319 | 0.244 | 0.172 | 0.511 | 0.429 | 1.078 | |
| ChatGPT-4o-mini2026.06 | 0.522 | 0.41 | 0.323 | 0.269 | 0.574 | 0.478 | 1.636 | |
| Watch-and-Act2026.06 | 0.394 | 0.271 | 0.248 | 0.187 | 0.405 | 0.311 | 1.434 | |
| DeepseekVL7B2026.06 | 0.368 | 0.275 | 0.227 | 0.182 | 0.384 | 0.334 | 0.801 | |
| V2CNet2026.06 | 0.357 | 0.231 | 0.201 | 0.153 | 0.369 | 0.264 | 0.895 | |
| Video2Command2026.06 | 0.312 | 0.196 | 0.173 | 0.15 | 0.33 | 0.234 | 1.103 | |
| LLaVA-NeXT7B2026.06 | 0.28 | 0.171 | 0.105 | 0.075 | 0.347 | 0.234 | 0.493 |