Video Action Classification on Diving-48
92.7Top-1 AccMOSS-L
Evaluation Results
| Method | Links | |
|---|---|---|
| MOSS-L2026.04 | 92.7 | |
| MOSS-B2026.04 | 91.2 | |
| Video-FocalNet-B2026.04 | 90.8 | |
| AIM ViT-L2026.04 | 90.6 | |
| V-JEPA 2 ViT-g384Parameters=1B, Input Resolution=384x384, Input format=32x4x3 frames2025.06 | 90.2 | |
| V-JEPA 2 ViT-gParameters=1B, Input Resolution=256x256, Input format=32x4x3 frames2025.06 | 90.1 | |
| V-JEPA 2 ViT-HParameters=600M, Input Resolution=256x256, Input format=32x4x3 frames2025.06 | 89.8 | |
| V-JEPA 22026.04 | 89.8 | |
| V-JEPA 2 ViT-LParameters=300M, Input Resolution=256x256, Input format=32x4x3 frames2025.06 | 89 | |
| Side4Video-B*2026.04 | 88.6 | |
| StructViT-B-4-12026.04 | 88.3 | |
| ORVIT TimeSformerPretrain=IN-21K, Frames=32, uses_bounding_boxes=true2021.10 | 88 | |
| ORViT2026.04 | 88 | |
| V-JEPA ViT-HParameters=600M, Evaluation Protocol=Attentive probe, Input Resolution=256x2562025.06 | 87.9 | |
| V-JEPA2026.04 | 87.9 | |
| InternVideo2_s2-1BParameters=1B, Evaluation Protocol=Attentive probe, Input Resolution=256x2562025.06 | 86.4 | |
| TimeSformer + STRG + STINPretrain=IN-21K, Frames=32, uses_bounding_boxes=true2021.10 | 83.5 | |
| DINOv2Parameters=1.1B, Evaluation Protocol=Attentive probe, Input Resolution=256x2562025.06 | 82.5 | |
| TQN†Pretrain=K400, Frames=ALL, uses_bounding_boxes=false2021.10 | 81.8 | |
| TimeSformer-L†Pretrain=IN-21K, Frames=96, uses_bounding_boxes=false2021.10 | 81 | |
| TimeSformer + STINPretrain=IN-21K, Frames=32, uses_bounding_boxes=true2021.10 | 81 | |
| TimeSformer-L2026.04 | 81 | |
| TimeSformer†Pretrain=IN-21K, Frames=32, uses_bounding_boxes=false2021.10 | 80 | |
| TimeSformer + STRGPretrain=IN-21K, Frames=32, uses_bounding_boxes=true2021.10 | 78.1 | |
| TimeSformer-HR2026.04 | 78 | |
| SlowFast, R101*Pretrain=K400, Frames=16, uses_bounding_boxes=false2021.10 | 77.6 | |
| PEcoreGParameters=1.9B, Evaluation Protocol=Attentive probe, Input Resolution=256x2562025.06 | 76.9 | |
| SigLIP2Parameters=1.2B, Evaluation Protocol=Attentive probe, Input Resolution=256x2562025.06 | 75.3 | |
| TimeSformer†Pretrain=IN-21K, Frames=16, uses_bounding_boxes=false2021.10 | 74.9 | |
| VideoPrismParameters=1B, Evaluation Protocol=Reported in Literature2025.06 | 71.3 | |
| OV-Encoder (Codec)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 69.4 | |
| OV-Encoder (Codec)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 67.2 | |
| OV-Encoder (Frame)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 63.2 | |
| DINOv3Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 61.3 | |
| DINOv3Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 58.6 | |
| OV-Encoder (Frame)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 57.6 | |
| SigLIP2Backbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 56.7 | |
| AIMv2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 55.7 | |
| SigLIPBackbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 54.7 | |
| SigLIP2Backbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 50.1 | |
| MetaCLIP2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 48 | |
| CLIPBackbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 46.6 | |
| CorrNet-101Pretrain=Sports1M, Two stream=False2019.06 | 44.7 | |
| CorrNet-101Pretrain=Sports1M, Two stream=false, test_clips=302019.06 | 44.7 | |
| SigLIPBackbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 43.9 | |
| AIMv2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 43.6 | |
| MetaCLIP2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 42.1 | |
| GST-50Pretrain=ImageNet, Two stream=False2019.06 | 38.8 | |
| GST-50Pretrain=ImageNet, Two stream=false2019.06 | 38.8 | |
| CorrNet-101Pretrain=None, Two stream=False, Note=Multi-clip sampling2019.06 | 38.6 | |
| CorrNet-101Pretrain=None, Two stream=false, test_clips=302019.06 | 38.6 | |
| CorrNet-101Pretrain=None, Two stream=False2019.06 | 38.2 | |
| CorrNet-101Pretrain=None, Two stream=false, test_clips=102019.06 | 38.2 | |
| CorrNet-50Pretrain=None, Two stream=False2019.06 | 37.9 | |
| CorrNet-50Pretrain=None, Two stream=false2019.06 | 37.9 | |
| DiMoFsPretrain=Kinetics, Two stream=False2019.06 | 31.4 | |
| DiMoFsPretrain=Kinetics, Two stream=false2019.06 | 31.4 | |
| R(2+1)DPretrain=Sports1M, Two stream=False2019.06 | 28.9 | |
| R(2+1)DPretrain=Sports1M, Two stream=false2019.06 | 28.9 | |
| MetaCLIPBackbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 28.9 | |
| TRNPretrain=ImageNet, Two stream=True2019.06 | 22.8 | |
| TRNPretrain=ImageNet, Two stream=true2019.06 | 22.8 | |
| R(2+1)DPretrain=None, Two stream=False2019.06 | 21.4 | |
| R(2+1)DPretrain=None, Two stream=false2019.06 | 21.4 |