Action Recognition on UCF-101 (3-Fold Cross-Validation)
99.63-Fold AccuracyOmniVec
Evaluation Results
| Method | Links | |
|---|---|---|
| OmniVecModality=Video2025.07 | 99.6 | |
| OmniVec2Modality=Video2025.07 | 99.1 | |
| CoViAR + flowInput Streams=Compressed Representation + Optical Flow2017.12 | 94.9 | |
| Very deep two-streamYear=20152015.07 | 91.4 | |
| CoViARInput Streams=Compressed Representation2017.12 | 90.4 | |
| TDD+FVYear=20152015.07 | 90.3 | |
| MIFS+FVYear=20152015.07 | 89.1 | |
| FlowInput Streams=Optical Flow2017.12 | 89 | |
| LSTMInputs=Optical Flow + Image Frames, Unroll Frames=302015.03 | 88.6 | |
| Two-stream+LSTMYear=20152015.07 | 88.6 | |
| Conv PoolingInputs=Image Frames + Optical Flow, Frames=1202015.03 | 88.2 | |
| Two-stream modelFusion Method=SVM2014.06 | 88 | |
| Two-Stream CNNInputs=Optical Flow + Image Frames, Fusion Method=SVM Fusion2015.03 | 88 | |
| Two-streamYear=20142015.07 | 88 | |
| IDT with higher-dimensional encodings2014.06 | 87.9 | |
| Improved Dense Trajectories (IDTF)s2015.03 | 87.9 | |
| iDT+HSVYear=20142015.07 | 87.9 | |
| Conv PoolingInputs=Image Frames + Optical Flow, Frames=302015.03 | 87.6 | |
| Two-stream modelFusion Method=averaging2014.06 | 86.9 | |
| Two-Stream CNNInputs=Optical Flow + Image Frames, Fusion Method=Averaging2015.03 | 86.9 | |
| Improved dense trajectories (IDT)2014.06 | 85.9 | |
| iDT+FVYear=20132015.07 | 85.9 | |
| Temporal stream ConvNetStream=Temporal2014.06 | 83.7 | |
| Single Frame CNN ModelInputs=Optical Flow2015.03 | 73.9 | |
| Single Frame Model2015.03 | 73.3 | |
| Spatial stream ConvNetStream=Spatial2014.06 | 73 | |
| Single Frame CNN ModelInputs=Images2015.03 | 73 | |
| Slow fusion spatio-temporal ConvNet2014.06 | 65.4 | |
| Slow Fusion CNN2015.03 | 65.4 | |
| DeepNetYear=20142015.07 | 63.3 | |
| Tuple verificationInitialization=Tuple verification2016.03 | 50.2 | |
| RandomInitialization=Random2016.03 | 38.6 |