Video Recognition on Kinetics-400 close-set
87.2Top-1 AccMoTE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MoTEEncoder=ViT-L/14, Input Size=16×224, GFLOPs × Views=1299×4×3, Param (M)=346.6, Unified model=true2024.10 | 87.2 | 97.7 | |
| DISTEncoder=ViT-L/14, Input Size=8×224, GFLOPs × Views=710×3×1, Param (M)=343.0, Unified model=true2024.10 | 86.9 | 97.6 | |
| AIMEncoder=ViT-L/14, Input Size=8×224, GFLOPs × Views=934×3×1, Param (M)=341.02024.10 | 86.8 | 97.2 | |
| MoTEEncoder=ViT-L/14, Input Size=8×224, GFLOPs × Views=649×4×3, Param (M)=346.6, Unified model=true2024.10 | 86.8 | 97.5 | |
| Text4VisEncoder=ViT-L/14, Input Size=8×224, GFLOPs × Views=649×4×3, Param (M)=346.6, Unified model=true2024.10 | 86.7 | 97.4 | |
| X-FlorenceEncoder=Florence, Input Size=32×224, GFLOPs × Views=2822×4×3, Unified model=false2024.10 | 86.5 | 96.9 | |
| ViFi-CLIPEncoder=ViT-B/16, Input Size=16×224, GFLOPs × Views=281×4×3, Param (M)=149.6, Unified model=false2024.10 | 83.9 | 96.3 | |
| Open-VCLIPEncoder=ViT-L/14, Input Size=8×224, GFLOPs × Views=/x3x1, Param (M)=304.0, Unified model=true2024.10 | 83.9 | 96.5 | |
| X-CLIPEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=145×4×3, Param (M)=131.5, Unified model=false2024.10 | 83.8 | 96.7 | |
| MoTEEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=141×4×3, Param (M)=98.8, Unified model=true2024.10 | 83 | 96.3 | |
| Text4VisEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=141×4×3, Param (M)=105.2, Unified model=true2024.10 | 82.9 | — | |
| ActionCLIPEncoder=ViT-B/16, Input Size=16×224, GFLOPs × Views=282×10×3, Param (M)=141.7, Unified model=false2024.10 | 82.6 | 96.2 | |
| X-CLIPEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=145×1×1, Param (M)=131.5, Unified model=false2024.10 | 82.3 | 95.8 | |
| ST-AdapterEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=152×3×1, Param (M)=93.0, Unified model=true2024.10 | 82 | 95.7 | |
| MoTEEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=141×1×1, Param (M)=98.8, Unified model=true2024.10 | 81.8 | 95.9 | |
| Vita-CLIPEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=/x4x3, Param (M)=125.0, Unified model=true2024.10 | 81.8 | 95.9 | |
| Vita-CLIPEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=/x1x1, Param (M)=125.0, Unified model=true2024.10 | 80.5 | 95.9 | |
| Open-VCLIPEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=/x1x1, Param (M)=86.2, Unified model=true2024.10 | 78.9 | — | |
| FROSTEREncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=/x1x1, Param (M)=86.2, Unified model=true2024.10 | 78.9 | 94.8 | |
| SiFEncoder=ViT-B/16, Input Size=8×224, GFLOPs × Views=142×4×3, Param (M)=143.92024.10 | 77.4 | 93.6 | |
| A6Encoder=ViT-B/16, Input Size=16×224, Param (M)=91.22024.10 | 76.9 | 93.5 |