Frame Selection for long-form video QA on 10-minute video 600 frames at 1 FPS, K=16
0.1E2E Latency (s)FFS
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FFSTrain-Free=false, Query Representation=Learned policy, Selection Evidence=Vis.2026.03 | 0.1 | 0.1 | — | |
| VidF4Train-Free=false, Query Representation=Learned score fn., Selection Evidence=Vis.2026.03 | 1.1 | 0.7 | — | |
| Frame-VoyagerTrain-Free=false, Query Representation=Learned score fn., Selection Evidence=Vis.2026.03 | 1.6 | 0.2 | — | |
| MDP3Train-Free=true, Query Representation=Global embedding, Selection Evidence=Vis.2026.03 | 1.8 | 0.7 | — | |
| AKSTrain-Free=true, Query Representation=Global embedding, Selection Evidence=Vis.2026.03 | 2.7 | 1.4 | — | |
| BOLTTrain-Free=true, Query Representation=Global embedding, Selection Evidence=Vis.2026.03 | 3 | 0.3 | — | |
| NeuS-QATrain-Free=true, Query Representation=Temporal logic spec., Selection Evidence=Vis.2026.03 | 6.7 | 6.7 | — | |
| T*Train-Free=true, Query Representation=Flat object queries, Selection Evidence=Vis. (OVD)2026.03 | 13 | 13 | — | |
| VSLSTrain-Free=true, Query Representation=Fixed relation triplets, Selection Evidence=Vis. (OVD)2026.03 | 13.3 | 13.3 | — | |
| HiMuTrain-Free=true, Query Representation=Hierarchical logic tree, Selection Evidence=Vis.+Aud.2026.03 | 13.3 | 9 | — | |
| VideoZoomerTrain-Free=false, Query Representation=Implicit (iterative VLM), Selection Evidence=Vis.2026.03 | 16 | 16 | — | |
| VideoAgentTrain-Free=false, Query Representation=Implicit (iterative LLM), Selection Evidence=Vis.2026.03 | 44.7 | 12.3 | — | |
| SeViLATrain-Free=false, Query Representation=Implicit (per-frame VLM), Selection Evidence=Vis.2026.03 | 60 | 60 | — |