Action Recognition on HMDB-51 (Base/Novel/HM Metrics)
92.3Base AccuracyVideo-STAR-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Video-STAR-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 92.3 | 91.9 | 92.1 | |
| Video-STAR-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 92.1 | 91.7 | 91.9 | |
| Motion-Guided Semantic Alignment with Negative PromptsBackbone=ViT-B/16 CLIP2026.04 | 78.5 | 60.4 | 68.3 | |
| VTD-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 78.4 | 63.5 | 70 | |
| GA2-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 78.3 | 58.9 | 67.2 | |
| ViFi-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 77.1 | 54.9 | 64.1 | |
| EZ-CLIPbackbone=ViT-162023.12 | 77 | 58.2 | 66.3 | |
| ViLT-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 76.7 | 57.5 | 65.7 | |
| SimVA2026.05 | 75.4 | 57.6 | 65.3 | |
| ZARBackbone=B/162026.04 | 75.2 | 55.2 | 63.7 | |
| AP-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 74.6 | 55.9 | 63.9 | |
| BDC-CLIP2026.05 | 74.5 | 55 | 63.3 | |
| FROSTERBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 74.1 | 58 | 65.1 | |
| ViFi-CLIPAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 73.8 | 53.3 | 61.9 | |
| ViFi CLIPSource=Rasheed et al. (2023)2023.12 | 73.8 | 53.3 | 61.9 | |
| ViFi-CLIPAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 73.8 | 53.3 | 61.9 | |
| ViFi-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 73.8 | 53.3 | 61.9 | |
| ViFi-CLIPBackbone=B/162026.04 | 73.8 | 53.3 | 61.9 | |
| ViFi-CLIP2026.05 | 73.8 | 53.3 | 61.9 | |
| Open-MeDeBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 73.6 | 56.4 | 63.9 | |
| TC-CLIP2026.05 | 73.3 | 59.1 | 65.5 | |
| CLIP text-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 70 | 51.2 | 59.1 | |
| CLIP text-FTAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 70 | 51.2 | 59.1 | |
| XCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 69.4 | 45.5 | 55 | |
| XCLIPSource=Ni et al. (2022)2023.12 | 69.4 | 45.5 | 55 | |
| X-CLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 69.4 | 45.5 | 55 | |
| X-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 69.4 | 45.5 | 55 | |
| XCLIPBackbone=B/162026.04 | 69.4 | 45.5 | 55 | |
| X-CLIP2026.05 | 69.4 | 45.5 | 55 | |
| ActionCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 69.1 | 37.3 | 48.5 | |
| ActionCLIPSource=Wang et al. (2021)2023.12 | 69.1 | 37.3 | 48.5 | |
| ActionCLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 69.1 | 37.3 | 48.5 | |
| ActionCLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 69.1 | 37.3 | 48.5 | |
| ActionCLIPBackbone=B/162026.04 | 69.1 | 37.3 | 48.5 | |
| ActionCLIP2026.05 | 69.1 | 37.3 | 48.5 | |
| ST-AdapterBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 65.3 | 48.9 | 55.9 | |
| CLIP image-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 62.6 | 47.5 | 54 | |
| CLIP image-FTAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 62.6 | 47.5 | 54 | |
| Qwen2.5-VL-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 54 | 40.8 | 46.5 | |
| Vanilla CLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 53.3 | 46.8 | 49.8 | |
| Vanila CLIPSource=Radford et al. (2021)2023.12 | 53.3 | 46.8 | 49.8 | |
| Vanilla CLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 53.3 | 46.8 | 49.8 | |
| Vanilla CLIPBackbone=B/162026.04 | 53.3 | 46.8 | 49.8 | |
| CLIP2026.05 | 53.3 | 46.8 | 49.8 | |
| A5Adaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 46.2 | 16 | 23.8 | |
| A5Source=Ju et al. (2022)2023.12 | 46.2 | 16 | 23.8 | |
| A5Adaptation Strategy=Adapting pre-trained image VL models2025.11 | 46.2 | 16 | 23.8 | |
| VPTBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 46.2 | 16 | 23.8 | |
| A52026.04 | 46.2 | 16 | 23.8 | |
| A52026.05 | 46.2 | 16 | 23.8 | |
| Qwen2.5-VL-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 41.7 | 50.3 | 45.6 |