Action Recognition on Epic Kitchens 100
48Top-1 AccOmniMAE
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| OmniMAEArch.=ViT-H, Pretrain Data=IN1K + SSv22022.06 | 48 | — | — | — | — | — | — | |
| MoViNet-A6GFLOPS=117, PARAM=31.4M, Frames=32, FPS=122021.03 | 47.7 | — | — | — | — | — | — | |
| OmniMAEArch.=ViT-L, Pretrain Data=IN1K + SSv22022.06 | 45.1 | — | — | — | — | — | — | |
| MoViNet-A5GFLOPS=74.9, PARAM=15.7M, Frames=32, FPS=122021.03 | 44.5 | — | — | — | — | — | — | |
| MoViNet-A4GFLOPS=42.2, PARAM=4.9M, Frames=32, FPS=122021.03 | 44.4 | — | — | — | — | — | — | |
| ViViT-L/16x2GFLOPS=3410, PARAM=100M2021.03 | 44 | — | — | — | — | — | — | |
| MoViNet-A2GFLOPS=7.59, PARAM=4.8M, Frames=32, FPS=122021.03 | 41.2 | — | — | — | — | — | — | |
| OmniMAEArch.=ViT-B, Pretrain Data=IN1K + K4002022.06 | 40.3 | — | — | — | — | — | — | |
| OmniMAEArch.=ViT-B, Pretrain Data=IN1K + SSv22022.06 | 39.3 | — | — | — | — | — | — | |
| SlowFast2021.03 | 38.5 | — | — | — | — | — | — | |
| TSM2021.03 | 38.3 | — | — | — | — | — | — | |
| MoViNet-A0GFLOPS=1.74, PARAM=3.1M, Frames=32, FPS=122021.03 | 36.8 | — | — | — | — | — | — | |
| SMILEBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 34.4 | — | — | — | — | — | — | |
| SIGMABackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 34.2 | — | — | — | — | — | — | |
| TSN2021.03 | 33.2 | — | — | — | — | — | — | |
| VideoMAEBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 33.2 | — | — | — | — | — | — | |
| MGMAEBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 33.2 | — | — | — | — | — | — | |
| MGMBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 32.4 | — | — | — | — | — | — | |
| SMILE w/o motionBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 32.3 | — | — | — | — | — | — | |
| MMEBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 32.2 | — | — | — | — | — | — | |
| EVERESTBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 30.5 | — | — | — | — | — | — | |
| MVDBackbone=ViT-B, Pre-training Dataset=K400, Evaluation Protocol=Linear Probing2025.04 | 29.7 | — | — | — | — | — | — | |
| PlayItBackX3GFLOPS=122.82022.10 | 15.9 | 29.2 | — | — | — | — | — | |
| Slow-FastGFLOPS=35.12022.10 | 15.4 | 28.6 | — | — | — | — | — | |
| Damen et al.GFLOPS=N/A2022.10 | 14.5 | 28.2 | — | — | — | — | — | |
| MBT (A)GFLOPS=34.22022.10 | 13 | — | — | — | — | — | — | |
| AIMv2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 56.6 | 45.6 | |
| AIMv2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 58.3 | 46.2 | |
| Avion (ViT-B)Extra Pre-training Data=WIT + Ego4D2024.11 | — | — | 49.1 | 70 | 59.8 | — | — | |
| Avion (ViT-L)Extra Pre-training Data=WIT + Ego4D2024.11 | — | — | 54.4 | 73 | 65.4 | — | — | |
| CLIPBackbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 52.8 | 36.1 | |
| DINOv3Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 62.5 | 51.7 | |
| DINOv3Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 63.2 | 51.9 | |
| IPL (I3D)Extra Pre-training Data=K4002024.11 | — | — | 41 | 68.6 | 51.2 | — | — | |
| LaViLa (TSF-B)Extra Pre-training Data=WIT + Ego4D2024.11 | — | — | 46.9 | 69 | 58.4 | — | — | |
| LaViLa (TSF-L)Extra Pre-training Data=WIT + Ego4D2024.11 | — | — | 51 | 72 | 62.9 | — | — | |
| LVMAE (ViT-B)Extra Pre-training Data=None2024.11 | — | — | 47.3 | 73.1 | 56.8 | — | — | |
| LVMAE (ViT-B)Extra Pre-training Data=Unlabeled K7102024.11 | — | — | 47 | 73 | 56.3 | — | — | |
| LVMAE (ViT-L)Extra Pre-training Data=Unlabeled K7102024.11 | — | — | 50.9 | 75.5 | 59.6 | — | — | |
| LVMAE (ViT-L)Extra Pre-training Data=K7102024.11 | — | — | 52.1 | 75 | 61.8 | — | — | |
| MeMViT-16, 16x4Extra Pre-training Data=K4002024.11 | — | — | 46.2 | 70.6 | 58.5 | — | — | |
| MeMViT-24, 32x3Extra Pre-training Data=K6002024.11 | — | — | 48.4 | 71.4 | 60.3 | — | — | |
| MetaCLIPBackbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 54.1 | 37.1 | |
| MetaCLIP2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 48 | 40.9 | |
| MetaCLIP2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 49.2 | 43.2 | |
| MoViNet-A5Extra Pre-training Data=N/A2024.11 | — | — | 44.5 | 69.1 | 55.1 | — | — | |
| MTV-BExtra Pre-training Data=IN21K2024.11 | — | — | 46.7 | 67.8 | 60.5 | — | — | |
| MTV-B↑280^2Extra Pre-training Data=WTS-60M2024.11 | — | — | 50.5 | 69.9 | 63.9 | — | — | |
| OmnivoreBackbone=Swin-T, Input Modality=RGB2025.12 | — | — | 35.9 | — | — | 62.8 | 47.8 | |
| Omnivore (Swin-B)Extra Pre-training Data=IN-(21K+1K)+K400+SUN2024.11 | — | — | 49.9 | 69.5 | 61.7 | — | — | |
| OV-Encoder (Codec)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 62.3 | 53.9 | |
| OV-Encoder (Codec)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 63.3 | 54.4 | |
| OV-Encoder (Frame)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 61.4 | 52.5 | |
| OV-Encoder (Frame)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 62.9 | 54.5 | |
| SigLIPBackbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 52.2 | 39.1 | |
| SigLIPBackbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 54.1 | 40.2 | |
| SigLIP2Backbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 54.2 | 43.8 | |
| SigLIP2Backbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | — | — | — | — | — | 56.4 | 45.2 | |
| SlowFastExtra Pre-training Data=K4002024.11 | — | — | 38.5 | 65.6 | 50 | — | — | |
| StudentBackbone=Swin-T, Input Modality=RGB, lambda=1, gamma=302025.12 | — | — | 39.3 | — | — | 65.4 | 51.7 | |
| TAdaFormer-B/16Extra Pre-training Data=K7102024.11 | — | — | 49.1 | 71 | 60.5 | — | — | |
| TAdaFormer-L/16Extra Pre-training Data=K7102024.11 | — | — | 51.8 | 71.7 | 64.1 | — | — | |
| ViViT-L/16x2Extra Pre-training Data=IN21K+K4002024.11 | — | — | 44 | 66.4 | 56.8 | — | — |