Sports Video Captioning on BFMD singles
47.1BLEU-1Ours
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| OursInput Modality=RGB-only2026.03 | 47.1 | 31.5 | 21.9 | 15.7 | 23.7 | 35.9 | 32.3 | |
| Shot2TacticModel Group=Vision-based Sports Captioning Models, Input Modality=RGB-only2026.03 | 45 | 30.1 | 20.8 | 14.6 | 22.8 | 34.9 | 27.9 | |
| InternVideo2Model Group=Pretrained Video-Language Models, Input Modality=RGB-only2026.03 | 42.4 | 27.4 | 18.4 | 12.7 | 22.8 | 33.4 | 23 | |
| Vid2SeqModel Group=Pretrained Video-Language Models, Input Modality=RGB-only2026.03 | 41.2 | 25.1 | 16.4 | 11.5 | 23.5 | 31.5 | 21.5 | |
| SoccerNet-CaptionModel Group=Vision-based Sports Captioning Models, Input Modality=RGB-only2026.03 | 38.4 | 24.8 | 16.2 | 10.7 | 20.1 | 32.8 | 11.9 | |
| GPT-5.2Model Group=Large Vision-Language Models, Evaluation Protocol=Zero-shot, Input Modality=RGB-only2026.03 | 24.2 | 12.4 | 6.4 | 3.7 | 16.2 | 21.4 | 6 | |
| GPT-4.1Model Group=Large Vision-Language Models, Evaluation Protocol=Zero-shot, Input Modality=RGB-only2026.03 | 22 | 12.9 | 7.8 | 5.2 | 16.9 | 23 | 6.8 | |
| Qwen2.5-VL-7B-InstructModel Group=Large Vision-Language Models, Evaluation Protocol=Zero-shot, Input Modality=RGB-only2026.03 | 18.8 | 6.5 | 2.2 | 1 | 14.6 | 15.5 | 2.2 | |
| Qwen3-VL-8B-InstructModel Group=Large Vision-Language Models, Evaluation Protocol=Zero-shot, Input Modality=RGB-only2026.03 | 17.2 | 8.2 | 3 | 1.3 | 14.8 | 16.7 | 2.4 |