Egocentric Video Understanding on EgoSchema
61.4ScoreAOT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| AOTPrefilling FLOPs (T)=7.0, FLOPs Ratio=17.2%, Before LLM Retained Ratio=20%2026.03 | 61.4 | — | — | — | |
| AOTPrefilling FLOPs (T)=5.2, FLOPs Ratio=12.7%, Before LLM Retained Ratio=15%2026.03 | 61.3 | — | — | — | |
| AOTPrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 61 | — | — | — | |
| LLaVA-OV-7BPrefilling FLOPs (T)=40.8, FLOPs Ratio=100%, Before LLM Retained Ratio=100%2026.03 | 60.4 | — | — | — | |
| VisionZipPrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 60.3 | — | — | — | |
| AOTPrefilling FLOPs (T)=3.4, FLOPs Ratio=8.3%, Before LLM Retained Ratio=10%2026.03 | 60.3 | — | — | — | |
| PruneVidPrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 59.9 | — | — | — | |
| VisionZipPrefilling FLOPs (T)=7.0, FLOPs Ratio=17.2%, Before LLM Retained Ratio=20%2026.03 | 59.8 | — | — | — | |
| VisionZipPrefilling FLOPs (T)=5.2, FLOPs Ratio=12.7%, Before LLM Retained Ratio=15%2026.03 | 59.8 | — | — | — | |
| PruneVidPrefilling FLOPs (T)=3.4, FLOPs Ratio=8.3%, Before LLM Retained Ratio=10%2026.03 | 59.8 | — | — | — | |
| PruneVidPrefilling FLOPs (T)=7.0, FLOPs Ratio=17.2%, Before LLM Retained Ratio=20%2026.03 | 59.7 | — | — | — | |
| PruneVidPrefilling FLOPs (T)=5.2, FLOPs Ratio=12.7%, Before LLM Retained Ratio=15%2026.03 | 59.7 | — | — | — | |
| DyCokePrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 59.5 | — | — | — | |
| PDropPrefilling FLOPs (T)=10.5, FLOPs Ratio=25.7%, Before LLM Retained Ratio=100%2026.03 | 58 | — | — | — | |
| VisionZipPrefilling FLOPs (T)=3.4, FLOPs Ratio=8.3%, Before LLM Retained Ratio=10%2026.03 | 58 | — | — | — | |
| FastVPrefilling FLOPs (T)=9.3, FLOPs Ratio=22.8%, Before LLM Retained Ratio=100%2026.03 | 57.5 | — | — | — | |
| LLaVA-Video-7BBackbone=LLaVA-Video-7B, FLOPs (T)=80.9, FLOPs Ratio=100%, Retention Ratio=100%2026.03 | 57.2 | — | — | — | |
| OursBackbone=LLaVA-Video-7B, FLOPs (T)=1.4, FLOPs Ratio=1.7%, Retention Ratio=2%/1%2026.03 | 46.8 | — | — | — | |
| Ours (w/o M)Backbone=LLaVA-Video-7B, FLOPs (T)=1.5, FLOPs Ratio=1.9%, Retention Ratio=2%2026.03 | 46.7 | — | — | — | |
| HoliTomBackbone=LLaVA-Video-7B, FLOPs (T)=1.4, FLOPs Ratio=1.7%, Retention Ratio=2%/1%2026.03 | 46.5 | — | — | — | |
| FastVidBackbone=LLaVA-Video-7B, FLOPs (T)=1.5, FLOPs Ratio=1.9%, Retention Ratio=2%2026.03 | 45.7 | — | — | — | |
| Ours (w/o M)Backbone=LLaVA-Video-7B, FLOPs (T)=1.2, FLOPs Ratio=1.5%, Retention Ratio=1%2026.03 | 44.8 | — | — | — | |
| OursBackbone=LLaVA-Video-7B, FLOPs (T)=1.1, FLOPs Ratio=1.4%, Retention Ratio=1%/0.5%2026.03 | 44.8 | — | — | — | |
| HoliTomBackbone=LLaVA-Video-7B, FLOPs (T)=1.1, FLOPs Ratio=1.4%, Retention Ratio=1%/0.5%2026.03 | 41.1 | — | — | — | |
| FsatVBackbone=LLaVA-Video-7B, FLOPs (T)=10.2, FLOPs Ratio=12.6%, Retention Ratio=100%/2%2026.03 | 41 | — | — | — | |
| VisionZipBackbone=LLaVA-Video-7B, FLOPs (T)=1.5, FLOPs Ratio=1.9%, Retention Ratio=2%2026.03 | 39.3 | — | — | — | |
| FsatVBackbone=LLaVA-Video-7B, FLOPs (T)=9.7, FLOPs Ratio=11.9%, Retention Ratio=100%/1%2026.03 | 38.8 | — | — | — | |
| FastVidBackbone=LLaVA-Video-7B, FLOPs (T)=1.2, FLOPs Ratio=1.5%, Retention Ratio=1%2026.03 | 38.8 | — | — | — | |
| VisionZipBackbone=LLaVA-Video-7B, FLOPs (T)=1.2, FLOPs Ratio=1.5%, Retention Ratio=1%2026.03 | 36.8 | — | — | — | |
| LLaVA-OV-0.5BBackbone=LLaVA-OV-0.5B, FLOPs (T)=3.54, FLOPs Ratio=100%, Retention Ratio=100%2026.03 | 26.6 | — | — | — | |
| OursBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.06, FLOPs Ratio=1.8%, Retention Ratio=2%/1%2026.03 | 24.8 | — | — | — | |
| Ours (w/o M)Backbone=LLaVA-OV-0.5B, FLOPs (T)=0.07, FLOPs Ratio=1.9%, Retention Ratio=2%2026.03 | 24.6 | — | — | — | |
| HoliTomBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.06, FLOPs Ratio=1.8%, Retention Ratio=2%/1%2026.03 | 24.5 | — | — | — | |
| OursBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.05, FLOPs Ratio=1.3%, Retention Ratio=1%/0.5%2026.03 | 24.2 | — | — | — | |
| Ours (w/o M)Backbone=LLaVA-OV-0.5B, FLOPs (T)=0.05, FLOPs Ratio=1.3%, Retention Ratio=1%2026.03 | 23.2 | — | — | — | |
| HoliTomBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.05, FLOPs Ratio=1.3%, Retention Ratio=1%/0.5%2026.03 | 23 | — | — | — | |
| FastVidBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.07, FLOPs Ratio=1.9%, Retention Ratio=2%2026.03 | 22.9 | — | — | — | |
| FastVBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.50, FLOPs Ratio=14.1%, Retention Ratio=100%/2%2026.03 | 21.2 | — | — | — | |
| FastVidBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.05, FLOPs Ratio=1.3%, Retention Ratio=1%2026.03 | 20.8 | — | — | — | |
| VisionZipBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.07, FLOPs Ratio=1.9%, Retention Ratio=2%2026.03 | 20 | — | — | — | |
| FastVBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.48, FLOPs Ratio=13.7%, Retention Ratio=100%/1%2026.03 | 19.3 | — | — | — | |
| VisionZipBackbone=LLaVA-OV-0.5B, FLOPs (T)=0.05, FLOPs Ratio=1.3%, Retention Ratio=1%2026.03 | 19.1 | — | — | — | |
| FastVBase Model=LLaVA-OneVision, Retention Ratio R=25%2026.02 | — | 60.4 | 57.8 | — | |
| FastVBase Model=LLaVA-OneVision, Retention Ratio R=20%2026.02 | — | 60.6 | 57.6 | — | |
| FastVBase Model=LLaVA-OneVision, Retention Ratio R=15%2026.02 | — | 59.8 | 56.6 | — | |
| FastVBase Model=LLaVA-OneVision, Retention Ratio R=10%2026.02 | — | 59 | 56 | — | |
| FastVBase Model=LLaVA-Video, Retention Ratio R=20%2026.02 | — | 54.8 | 54.1 | — | |
| FastVBase Model=LLaVA-Video, Retention Ratio R=10%2026.02 | — | 50.6 | 51.1 | — | |
| FastVRetention Ratio R=25%2026.02 | — | 60.2 | 57.2 | — | |
| FastVRetention Ratio R=15%2026.02 | — | 59.6 | 56.5 | — | |
| FastVIDBase Model=LLaVA-OneVision, Retention Ratio R=25%2026.02 | — | 61.2 | 59.5 | — | |
| FastVIDBase Model=LLaVA-OneVision, Retention Ratio R=20%2026.02 | — | 61.2 | 59.5 | — | |
| FastVIDBase Model=LLaVA-OneVision, Retention Ratio R=15%2026.02 | — | 58.8 | 58.9 | — | |
| FastVIDBase Model=LLaVA-OneVision, Retention Ratio R=10%2026.02 | — | 58.8 | 58.7 | — | |
| FastVIDBase Model=LLaVA-Video, Retention Ratio R=20%2026.02 | — | 57 | 55 | — | |
| FastVIDBase Model=LLaVA-Video, Retention Ratio R=10%2026.02 | — | 54.8 | 52.4 | — | |
| FastVIDRetention Ratio R=25%2026.02 | — | 58.2 | 56.7 | — | |
| FastVIDRetention Ratio R=15%2026.02 | — | 56.6 | 56 | — | |
| FlashVIDBase Model=LLaVA-OneVision, Retention Ratio R=25%2026.02 | — | 63.4 | 60.4 | — | |
| FlashVIDBase Model=LLaVA-OneVision, Retention Ratio R=20%2026.02 | — | 63 | 60.1 | — | |
| FlashVIDBase Model=LLaVA-OneVision, Retention Ratio R=15%2026.02 | — | 62.8 | 60.4 | — | |
| FlashVIDBase Model=LLaVA-OneVision, Retention Ratio R=10%2026.02 | — | 62.4 | 60 | — | |
| FlashVIDBase Model=LLaVA-Video, Retention Ratio R=20%2026.02 | — | 58.4 | 56.4 | — | |
| FlashVIDBase Model=LLaVA-Video, Retention Ratio R=10%2026.02 | — | 57.2 | 54.9 | — | |
| FlashVIDRetention Ratio R=25%2026.02 | — | 59.4 | 57.2 | — | |
| FlashVIDRetention Ratio R=15%2026.02 | — | 57.6 | 56.4 | — | |
| FrameOracleBackbone=Qwen2.5-VL-3B, Frames=32→20.92025.10 | — | — | — | 53.8 | |
| FrameOracleBackbone=Qwen2.5-VL-3B, Frames=128→27.82025.10 | — | — | — | 54.5 | |
| FrameOracleBackbone=LLaVA-OneVision-7B, Frames=16→10.42025.10 | — | — | — | 62.4 | |
| FrameOracleBackbone=LLaVA-OneVision-7B, Frames=64→13.92025.10 | — | — | — | 63.4 | |
| FrameOracleBackbone=LLaVA-Video-7B, Frames=16→10.42025.10 | — | — | — | 54.6 | |
| FrameOracleBackbone=LLaVA-Video-7B, Frames=64→13.92025.10 | — | — | — | 55.2 | |
| FrameOracleBackbone=VideoLLaMA3-7B, Frames=16→10.42025.10 | — | — | — | 61.8 | |
| FrameOracleBackbone=VideoLLaMA3-7B, Frames=64→13.92025.10 | — | — | — | 62.4 | |
| FrameOracleBackbone=Qwen3-VL-8B, Frames=32→20.92025.10 | — | — | — | 71.4 | |
| FrameOracleBackbone=Qwen3-VL-8B, Frames=128→27.82025.10 | — | — | — | 72.3 | |
| Gemini-1.5-Pro2026.05 | — | — | — | 72.2 | |
| Gemini-3-Flash2026.05 | — | — | — | 37 | |
| GPT-4o2026.05 | — | — | — | 72 | |
| InternVL2.5#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline2026.01 | — | — | — | 51.5 | |
| LLaDA-V#S (Model Size)=8B, #F (Number of input frames)=32, Architecture=DLM Baseline, reproduced=true2026.01 | — | — | — | 57.9 | |
| LLaVA-NeXT-Video#S (Model Size)=7B, #F (Number of input frames)=32, Architecture=AR Baseline2026.01 | — | — | — | 43.9 | |
| LLaVA-OneVision#S (Model Size)=7B, #F (Number of input frames)=32, Architecture=AR Baseline, reproduced=true2026.01 | — | — | — | 60.1 | |
| LLaVA-OneVision-7BFrames=162025.10 | — | — | — | 60.8 | |
| LLaVA-OV2026.05 | — | — | — | 60.1 | |
| LLaVA-Video#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline, reproduced=true2026.01 | — | — | — | 57.3 | |
| LLaVA-Video-7BFrames=162025.10 | — | — | — | 54.2 | |
| PruneVIDBase Model=LLaVA-OneVision, Retention Ratio R=25%2026.02 | — | 61 | 58.1 | — | |
| PruneVIDBase Model=LLaVA-OneVision, Retention Ratio R=20%2026.02 | — | 63.2 | 60.2 | — | |
| PruneVIDBase Model=LLaVA-OneVision, Retention Ratio R=15%2026.02 | — | 61.6 | 57.7 | — | |
| PruneVIDBase Model=LLaVA-OneVision, Retention Ratio R=10%2026.02 | — | 60 | 57.2 | — | |
| Qwen2-VL#S (Model Size)=7B, #F (Number of input frames)=2fps, Architecture=AR Baseline2026.01 | — | — | — | 66.7 | |
| Qwen2.5-VL#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline, reproduced=true2026.01 | — | — | — | 65 | |
| Qwen2.5-VL-3BFrames=322025.10 | — | — | — | 53.4 | |
| Qwen3-VL-8BFrames=322025.10 | — | — | — | 70.8 | |
| Qwen3.5-27B2026.05 | — | — | — | 22 | |
| STEMO-TrackBackbone=Gemini-3-Flash2026.05 | — | — | — | 78.4 | |
| STEMO-TrackBackbone=Qwen3-VL-235B2026.05 | — | — | — | 76.4 | |
| VanillaBase Model=LLaVA-OneVision, Retention Ratio R=100%2026.02 | — | 62.2 | 60.3 | — | |
| VanillaBase Model=LLaVA-Video, Retention Ratio R=100%2026.02 | — | 59.4 | 57.3 | — |