Multi-Instance Retrieval on EPIC-Kitchens 100 (test)
63.3mAP (Avg)EgoVideo
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| EgoVideo2024.06 | 63.3 | — | — | — | — | 67.6 | 58.9 | 75 | 71.5 | 73.2 | |
| AVION2024.06 | 54.5 | — | — | — | — | 57.9 | 51.1 | 70.4 | 67.6 | 69 | |
| LaViLA2024.06 | 50.9 | — | — | — | — | 54.7 | 47.1 | 68.1 | 64.9 | 66.5 | |
| TimeMambalearning_mode=fine-tuned, sampled_frames=162024.03 | 45.3 | — | — | — | — | 50.3 | 40.3 | 62.4 | 59.2 | 60.9 | |
| EgoVLPlearning_mode=fine-tuned, sampled_frames=162024.03 | 45 | — | — | — | — | 49.5 | 40.5 | 60.9 | 57.9 | 59.4 | |
| EgoVLP2024.06 | 45 | — | — | — | — | 49.9 | 40.5 | 60.9 | 57.9 | 59.4 | |
| TimeSformerlearning_mode=fine-tuned, sampled_frames=162024.03 | 44.2 | — | — | — | — | 49.1 | 39.3 | 60 | 57.6 | 58.8 | |
| TimeMambaAdaptation=Frozen [3], Mamba in Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=72024.03 | 26.8 | — | — | — | — | 30.7 | 22.8 | 31.3 | 27.8 | 29.5 | |
| TimeMambaAdaptation=Vanilla [6], Mamba in Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=62024.03 | 26.2 | — | — | — | — | 30.3 | 22.1 | 30.9 | 27.5 | 29.2 | |
| LaViLaAdaptation=Frozen [3], Attn in Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=32024.03 | 26 | — | — | — | — | — | — | — | — | 28.8 | |
| TimeSformerAdaptation=Frozen [3], Attn in Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=52024.03 | 26 | — | — | — | — | 29.8 | 22.2 | 30.6 | 27.5 | 29 | |
| TimeMambaAdaptation=Frozen [3], Mamba in Space-Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=82024.03 | 26 | — | — | — | — | 30.1 | 21.9 | 30.7 | 27.1 | 28.9 | |
| TimeSformerAdaptation=Vanilla [6], Attn in Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=42024.03 | 25.5 | — | — | — | — | 29.2 | 21.8 | 30.1 | 27.1 | 28.6 | |
| EgoVLPAdaptation=Frozen [3], Attn in Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=22024.03 | 23.3 | — | — | — | — | 26 | 20.6 | 28.8 | 27 | 27.9 | |
| EgoVLPAdaptation=Frozen [3], Attn in Time, Evaluation Protocol=Zero-shot, Frames Sampled=4, Reference ID=12024.03 | 16.6 | — | — | — | — | 19.4 | 13.9 | 24.1 | 22 | 23.1 | |
| AVIONTraining data regime=Original Narratives, Corpus size=4.0M, Hardware=8x A5000, Batch size=256, Memory=192023.09 | — | 28.4 | — | 130 | 11.06 | — | — | — | — | — | |
| AVIONTraining data regime=LLM-Augmented, Corpus size=35.0M, Hardware=8x A5000, Batch size=256, Memory=192023.09 | — | 32.7 | — | 260 | 22.12 | — | — | — | — | — | |
| EgoVLPEvaluation Protocol=Fine-tuned2023.01 | — | 45 | 59.4 | — | — | — | — | — | — | — | |
| EgoVLPTraining data regime=Original Narratives, Corpus size=3.8M, Hardware=32x A100, Batch size=16, Memory=222023.09 | — | 23.3 | — | 1,536 | 227.33 | — | — | — | — | — | |
| HierVL-AvgEvaluation Protocol=Fine-tuned2023.01 | — | 44.9 | 59.8 | — | — | — | — | — | — | — | |
| HierVL-SAEvaluation Protocol=Fine-tuned2023.01 | — | 46.7 | 61.1 | — | — | — | — | — | — | — | |
| HierVL-w/o HierEvaluation Protocol=Fine-tuned2023.01 | — | 44.7 | 59.8 | — | — | — | — | — | — | — | |
| JPOSEEvaluation Protocol=Fine-tuned, Backbone=TBN2023.01 | — | 44 | 53.5 | — | — | — | — | — | — | — | |
| LaViLaTraining data regime=LLM-Augmented, Corpus size=35.0M, Hardware=32x V100, Batch size=32, Memory=252023.09 | — | 30.9 | — | 1,824 | 202.46 | — | — | — | — | — | |
| MI-MMEvaluation Protocol=Fine-tuned, Backbone=S3D2023.01 | — | 29.2 | 44.7 | — | — | — | — | — | — | — | |
| MMEEvaluation Protocol=Fine-tuned, Backbone=TBN2023.01 | — | 38.5 | 48.5 | — | — | — | — | — | — | — |