Action Recognition on Kinetics-400 (Base, Novel, HM metrics)
96.3Base AccuracyVideo-STAR-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Video-STAR-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 96.3 | 97.2 | 96.7 | |
| Qwen2.5-VL-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 87.8 | 84.8 | 86.3 | |
| Video-STAR-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 86 | 86.4 | 86.2 | |
| Motion-Guided Semantic Alignment with Negative PromptsBackbone=ViT-B/16 CLIP2026.04 | 78.8 | 60.6 | 68.5 | |
| VTD-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 78.5 | 63.5 | 70.1 | |
| FROSTERBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 77.8 | 64.3 | 70.4 | |
| ViLT-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 77.4 | 63 | 69.5 | |
| Open-MeDeBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 77.2 | 63.8 | 69.9 | |
| AP-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 77.2 | 64.1 | 70 | |
| GA2-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 77 | 63.3 | 69.5 | |
| ViFi-CLIPAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 76.4 | 61.1 | 67.9 | |
| ViFi CLIPSource=Rasheed et al. (2023)2023.12 | 76.4 | 61.1 | 67.9 | |
| ViFi-CLIPAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 76.4 | 61.1 | 67.9 | |
| ViFi-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 76.4 | 61.1 | 67.9 | |
| ViFi-CLIPBackbone=B/162026.04 | 76.4 | 61.1 | 67.9 | |
| ZARBackbone=B/162026.04 | 75.2 | 60.7 | 67.2 | |
| XCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 74.1 | 56.4 | 64 | |
| XCLIPSource=Ni et al. (2022)2023.12 | 74.1 | 56.4 | 64 | |
| X-CLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 74.1 | 56.4 | 64 | |
| X-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 74.1 | 56.4 | 64 | |
| XCLIPBackbone=B/162026.04 | 74.1 | 56.4 | 64 | |
| A52026.04 | 74.1 | 56.4 | 64 | |
| ST-AdapterBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 73.6 | 62 | 67.3 | |
| CLIP text-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 73.4 | 59.7 | 65.8 | |
| CLIP text-FTAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 73.4 | 59.7 | 65.8 | |
| EZ-CLIPbackbone=ViT-162023.12 | 73.1 | 60.6 | 66.3 | |
| CLIP image-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 72.9 | 58 | 64.6 | |
| CLIP image-FTAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 72.9 | 58 | 64.6 | |
| A5Adaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 69.7 | 37.6 | 48.8 | |
| A5Source=Ju et al. (2022)2023.12 | 69.7 | 37.6 | 48.8 | |
| A5Adaptation Strategy=Adapting pre-trained image VL models2025.11 | 69.7 | 37.6 | 48.8 | |
| VPTBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 69.7 | 37.6 | 48.8 | |
| ActionCLIPBackbone=B/162026.04 | 69 | 57.2 | 62.6 | |
| Vanilla CLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 62.3 | 53.4 | 57.5 | |
| Vanila CLIPSource=Radford et al. (2021)2023.12 | 62.3 | 53.4 | 57.5 | |
| Vanilla CLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 62.3 | 53.4 | 57.5 | |
| ActionCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 61 | 46.2 | 52.6 | |
| ActionCLIPSource=Wang et al. (2021)2023.12 | 61 | 46.2 | 52.6 | |
| ActionCLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 61 | 46.2 | 52.6 | |
| ActionCLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 61 | 46.2 | 52.6 | |
| Vanilla CLIPBackbone=B/162026.04 | 53.3 | 46.8 | 49.8 | |
| Qwen2.5-VL-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 48.3 | 40.5 | 44.1 |