Action Recognition on UCF-101 (Base, Novel, HM Metrics)
99.6Base AccuracyVideo-STAR-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Video-STAR-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 99.6 | 99.8 | 99.7 | |
| Video-STAR-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 96.9 | 98.9 | 97.9 | |
| GA2-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 96.8 | 75.2 | 84.6 | |
| VTD-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 95.5 | 73.7 | 83.2 | |
| SimVA2026.05 | 95.5 | 82 | 88.2 | |
| TC-CLIP2026.05 | 95.4 | 81.6 | 88 | |
| FROSTERBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 95.3 | 80 | 87 | |
| ViLT-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 95.2 | 70.5 | 81 | |
| ViFi-CLIPAdaptation Strategy=Prompt tuning pre-trained image VL models2025.11 | 95.1 | 74.1 | 83.6 | |
| Open-MeDeBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 94.9 | 78.5 | 85.9 | |
| BDC-CLIP2026.05 | 94.9 | 78.8 | 86.1 | |
| AP-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 94.8 | 77 | 84.8 | |
| EZ-CLIPbackbone=ViT-162023.12 | 94.4 | 77.9 | 85.4 | |
| ViFi-CLIPAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 92.9 | 67.7 | 78.3 | |
| ViFi CLIPSource=Rasheed et al. (2023)2023.12 | 92.9 | 67.7 | 78.3 | |
| ViFi-CLIPAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 92.9 | 67.7 | 78.3 | |
| ViFi-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 92.9 | 67.7 | 78.3 | |
| ViFi-CLIP2026.05 | 92.9 | 67.7 | 78.3 | |
| CLIP text-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 90.9 | 67.4 | 77.4 | |
| CLIP text-FTAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 90.9 | 67.4 | 78.3 | |
| A5Adaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 90.5 | 40.4 | 55.8 | |
| A5Source=Ju et al. (2022)2023.12 | 90.5 | 40.4 | 55.8 | |
| A5Adaptation Strategy=Adapting pre-trained image VL models2025.11 | 90.5 | 40.4 | 55.8 | |
| VPTBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 90.5 | 40.4 | 55.8 | |
| A52026.05 | 90.5 | 40.4 | 55.8 | |
| ActionCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 90.1 | 58.1 | 70.7 | |
| ActionCLIPSource=Wang et al. (2021)2023.12 | 90.1 | 58.1 | 70.7 | |
| ActionCLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 90.1 | 58.1 | 70.7 | |
| ActionCLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 90.1 | 58.1 | 70.7 | |
| ActionCLIP2026.05 | 90.1 | 58.1 | 70.7 | |
| XCLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 89.9 | 58.9 | 71.2 | |
| XCLIPSource=Ni et al. (2022)2023.12 | 89.9 | 58.9 | 71.2 | |
| X-CLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 89.9 | 58.9 | 71.2 | |
| X-CLIPBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 89.9 | 58.9 | 71.2 | |
| X-CLIP2026.05 | 89.9 | 58.9 | 71.2 | |
| CLIP image-FTAdaptation Protocol=Tuning pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 86.4 | 65.3 | 74.4 | |
| CLIP image-FTAdaptation Strategy=Tuning pre-trained image VL models2025.11 | 86.4 | 65.3 | 74.4 | |
| ST-AdapterBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 85.5 | 76.8 | 80.9 | |
| Qwen2.5-VL-7BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 85 | 82.4 | 83.7 | |
| Vanilla CLIPAdaptation Protocol=Adapting pre-trained image VL models, Backbone=ViT-B/16, Number of frames=322022.12 | 78.5 | 63.6 | 70.3 | |
| Vanila CLIPSource=Radford et al. (2021)2023.12 | 78.5 | 63.6 | 70.3 | |
| Vanilla CLIPAdaptation Strategy=Adapting pre-trained image VL models2025.11 | 78.5 | 63.6 | 70.3 | |
| CLIP2026.05 | 78.5 | 63.6 | 70.3 | |
| Qwen2.5-VL-3BBackbone=ViT-B/16, Evaluation Setting=base-to-novel2025.10 | 71.4 | 62.4 | 66.6 |