Action Recognition on HMDB51 (val)
62.6AccuracyVideoMAE
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VideoMAEBackbone=ViT-B, Number of Frames=16, Evaluation Protocol=fine-tuned2022.03 | 62.6 | — | — | |
| TACOEncoder=ViT-L/142026.06 | 59.9 | — | — | |
| MoTEEncoder=ViT-L/142026.06 | 56.3 | — | — | |
| OTIEncoder=ViT-L/142026.06 | 55.8 | — | — | |
| TACOEncoder=ViT-B/162026.06 | 54.6 | — | — | |
| FROSTER*Encoder=ViT-B/162026.06 | 54.5 | — | — | |
| Open-VCLIP*Encoder=ViT-B/162026.06 | 53.2 | — | — | |
| BIKEEncoder=ViT-L/142026.06 | 52.8 | — | — | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, BLIP verbs, vis.encoder=ViT-B/16, frames=16/322023.03 | 52.3 | — | — | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, BLIP verbs, vis.encoder=ViT-B/16, frames=162023.03 | 52.2 | — | — | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, vis.encoder=ViT-B/16, frames=16/322023.03 | 51.9 | — | — | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, vis.encoder=ViT-B/16, frames=162023.03 | 51.6 | — | — | |
| ViFi-CLIPgt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 51.3 | — | — | |
| ViFi-CLIP (re-eval)gt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=162023.03 | 50.9 | — | — | |
| MAXIgt=no, language=K400 dict., vis.encoder=ViT-B/16, frames=162023.03 | 50.5 | — | — | |
| ST-AdapterEncoder=ViT-B/162026.06 | 50.3 | — | — | |
| Text4VisEncoder=ViT-L/142026.06 | 49.8 | — | — | |
| AIMEncoder=ViT-B/162026.06 | 49.5 | — | — | |
| CLIPEncoder=ViT-B/162026.06 | 46.7 | — | — | |
| XCLIPgt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 44.6 | — | — | |
| A5gt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 44.3 | — | — | |
| ActionCLIPgt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 40.8 | — | — | |
| MoCo v3Backbone=ViT-B, Number of Frames=16, Evaluation Protocol=fine-tuned2022.03 | 39.2 | — | — | |
| JigsawNetgt=yes, language=Manual description, vis.encoder=R(2+1)D, frames=162023.03 | 38.7 | — | — | |
| CLIPgt=no, vis.encoder=ViT-B/16, frames=162023.03 | 38 | — | — | |
| ER-ZSARgt=yes, language=Manual description, vis.encoder=TSM, frames=162023.03 | 35.3 | — | — | |
| E2EEvaluation protocol=Zero-shot, Number of frames=32, Inference view=Single-view, Model architecture type=Uni-modal2025.11 | 32.7 | — | — | |
| from scratchBackbone=ViT-B, Number of Frames=16, Evaluation Protocol=supervised training2022.03 | 18 | — | — | |
| Baseline (Naive + Var.)Data Type=3D Fractals2026.02 | — | 49.3 | 79.4 | |
| Data-Driven (RF-Filter)Data Type=3D Fractals2026.02 | — | 44.3 | 73.5 | |
| EVERESTTarget=Pixel, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 30.3 | — | |
| From ScratchData Type=N/A2026.02 | — | 31.5 | — | |
| Kinetics (Kay et al., 2017)Data Type=Natural images2026.02 | — | 70.1 | — | |
| MGMTarget=Pixel, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 40.3 | — | |
| MGMAETarget=Pixel, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 41.3 | — | |
| MMETarget=HOG, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 37.1 | — | |
| MVDTarget=Pixel, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 28.6 | — | |
| SIGMATarget=DINO, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 52.3 | — | |
| SMILETarget=CLIP, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 53.4 | — | |
| SVD-Control FilterData Type=3D Fractals2026.02 | — | 47.4 | 78.3 | |
| Svyezhentsev et al. (2024)Data Type=2D Fractals2026.02 | — | 56.5 | — | |
| Targeted Smart Filtering (TSF)Data Type=3D Fractals2026.02 | — | 48.5 | 79.5 | |
| TrackMAETarget=Pixel, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 40.6 | — | |
| TrackMAETarget=CLIP, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 53.1 | — | |
| VideoMAETarget=Pixel, Evaluation Protocol=Linear probing, Backbone=ViT-B, Pre-trained Dataset=K4002026.03 | — | 37.7 | — |