Action Recognition on Diving48 (test)
35.5Top-1 AccDeCoDe (Qwen2.5-VL-7B)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeCoDe (Qwen2.5-VL-7B)Setting=1-shot dec., Label configuration=With labels (Semantic)2026.06 | 35.5 | — | |
| V-JEPA 2Model Strategy=V-JEPA 2, Evaluation Protocol=Video Probing2026.05 | 30.08 | — | |
| TIME + V-JEPA 2Model Strategy=TIME + V-JEPA 2, Evaluation Protocol=Video Probing2026.05 | 30.04 | 0.04 | |
| TIME + VideoMAEv2Model Strategy=TIME + VideoMAEv2, Evaluation Protocol=Video Probing2026.05 | 24.97 | 0.73 | |
| Qwen2.5-VL-7BSetting=1-shot, Label configuration=With labels (Semantic)2026.06 | 24.7 | — | |
| VideoMAEv2Model Strategy=VideoMAEv2, Evaluation Protocol=Video Probing2026.05 | 24.24 | — | |
| Qwen2.5-VL-7BSetting=0-shot, Label configuration=With labels (Semantic)2026.06 | 22.7 | — | |
| TIME + DINOv3Model Strategy=TIME + DINOv3 (4f), Number of frames=4f, Evaluation Protocol=Video Probing2026.05 | 18.08 | 1.73 | |
| DINOv3Model Strategy=DINOv3 (4f), Number of frames=4f, Evaluation Protocol=Video Probing2026.05 | 16.35 | — | |
| TIME + CLIPModel Strategy=TIME + CLIP (4f), Number of frames=4f, Evaluation Protocol=Video Probing2026.05 | 16.11 | 2.14 | |
| CLIPModel Strategy=CLIP (4f), Number of frames=4f, Evaluation Protocol=Video Probing2026.05 | 13.97 | — | |
| TIME + RVMModel Strategy=TIME + RVM, Evaluation Protocol=Video Probing2026.05 | 11.74 | 2.53 | |
| RVMModel Strategy=RVM, Evaluation Protocol=Video Probing2026.05 | 9.21 | — | |
| TIMEModel Strategy=TIME (Ours), Evaluation Protocol=Video Probing2026.05 | 8.37 | — |