Action Recognition on Something-Something V2 (Base, Novel, HM)
19.2Base ScoreVideo-STAR-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Video-STAR-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 19.2 | 13 | 15.5 | |
| FROSTERBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 18.3 | 12.2 | 14.6 | |
| VTD-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 17.8 | 13.9 | 15.4 | |
| Open-MeDeBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 17.1 | 12.3 | 14.3 | |
| EZ-CLIPbackbone=ViT-162023.12 | 16.6 | 13.3 | 14.8 | |
| AP-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 16.3 | 12.9 | 14.4 | |
| ViFi-CLIPAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 16.2 | 12.1 | 13.9 | |
| ViFi CLIPSource=Rasheed et al. (2023)2023.12 | 16.2 | 12.1 | 13.9 | |
| ViFi-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 16.2 | 12.1 | 13.9 | |
| Qwen2.5-VL-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 14 | 9.9 | 11.6 | |
| Video-STAR-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 13.5 | 11.3 | 12.3 | |
| ActionCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 13.3 | 10.1 | 11.5 | |
| ActionCLIPSource=Wang et al. (2021)2023.12 | 13.3 | 10.1 | 11.5 | |
| ActionCLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 13.3 | 10.1 | 11.5 | |
| CLIP text-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 12.4 | 9.5 | 10.8 | |
| ST-AdapterBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 9.3 | 8.4 | 8.8 | |
| CLIP image-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 9.2 | 8.5 | 8.8 | |
| XCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 8.5 | 6.6 | 7.4 | |
| XCLIPSource=Ni et al. (2022)2023.12 | 8.5 | 6.6 | 7.4 | |
| X-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 8.5 | 6.6 | 7.4 | |
| A5Adaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 8.3 | 5.3 | 6.4 | |
| A5Source=Ju et al. (2022)2023.12 | 8.3 | 5.3 | 6.4 | |
| VPTBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 8.3 | 5.3 | 6.4 | |
| Vanilla CLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 4.9 | 5.3 | 5.1 | |
| Vanila CLIPSource=Radford et al. (2021)2023.12 | 4.9 | 5.3 | 5.1 | |
| Qwen2.5-VL-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 3.5 | 3.2 | 3.3 |