Temporal Question Grounding on NExT-GQA
0.442mIoUVideoAuto-R1
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| VideoAuto-R1Backbone=Qwen3-VL-8B2026.01 | 0.442 | — | — | 81.1 | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=6.8%, Reasoning=false2026.03 | 0.439 | — | — | 73.2 | |
| VITAL2026.01 | 0.43 | — | — | 78.7 | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=6.3%, Reasoning=false2026.03 | 0.391 | — | — | 73 | |
| Temporal-RLT2026.01 | 0.373 | — | — | 78.7 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B2026.01 | 0.367 | — | — | 80.6 | |
| VanillaBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=100%, Reasoning=false2026.03 | 0.366 | — | — | 81.1 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=16.1%, Reasoning=true2026.03 | 0.353 | — | — | 79.3 | |
| VanillaBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=100%, Reasoning=false2026.03 | 0.342 | — | — | 78.7 | |
| ToMeBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.34 | — | — | 79.2 | |
| FlashVidBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=31.3%, Reasoning=false2026.03 | 0.339 | — | — | 77.8 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=100%, Reasoning=true2026.03 | 0.338 | — | — | 73.6 | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=16.1%, Reasoning=false2026.03 | 0.333 | — | — | 76.8 | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=12.3%, Reasoning=false2026.03 | 0.326 | — | — | 75.4 | |
| FlashVidBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=12.6%, Reasoning=false2026.03 | 0.318 | — | — | 75.6 | |
| ToMeBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.317 | — | — | 77.1 | |
| ToMeBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.315 | — | — | 77.4 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=100%, Reasoning=true2026.03 | 0.31 | — | — | 68 | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=16.2%, Reasoning=false2026.03 | 0.302 | — | — | 75.1 | |
| VanillaBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=100%, Reasoning=false2026.03 | 0.299 | — | — | 79.8 | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=12.3%, Reasoning=false2026.03 | 0.299 | — | — | 74.2 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=6.8%, Reasoning=true2026.03 | 0.294 | — | — | 76.6 | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=6.8%, Reasoning=false2026.03 | 0.282 | — | — | 71.8 | |
| VanillaBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=100%, Reasoning=false2026.03 | 0.28 | — | — | 78.9 | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=6.3%, Reasoning=false2026.03 | 0.28 | — | — | 71.5 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=16.1%, Reasoning=false2026.03 | 0.272 | — | — | 78.1 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=6.8%, Reasoning=true2026.03 | 0.247 | — | — | 74.7 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=6.8%, Reasoning=false2026.03 | 0.239 | — | — | 76.2 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=16.2%, Reasoning=false2026.03 | 0.232 | — | — | 76.6 | |
| Random DropBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.224 | — | — | 79.3 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=6.8%, Reasoning=false2026.03 | 0.204 | — | — | 74.3 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, reproduced=true2026.01 | 0.202 | — | — | 53.3 | |
| TGBVision Encoder=OF+CNN2024.02 | 0.199 | 0.233 | 0.112 | — | |
| Random DropBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.199 | — | — | 76.9 | |
| TimeChat2026.01 | 0.174 | — | — | 28.8 | |
| Random DropBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.166 | — | — | 77.5 | |
| FlashVidBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=31.3%, Reasoning=false2026.03 | 0.165 | — | — | 78.1 | |
| ToMeBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.163 | — | — | 77.8 | |
| FlashVidBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=12.6%, Reasoning=false2026.03 | 0.161 | — | — | 77.4 | |
| ToMeBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.157 | — | — | 77.3 | |
| Random DropBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.156 | — | — | 77.2 | |
| Random DropBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.154 | — | — | 76.3 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=6.3%, Reasoning=false2026.03 | 0.154 | — | — | 74.1 | |
| IGVVision Encoder=ResNet2024.02 | 0.14 | 0.198 | 0.096 | — | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=12.3%, Reasoning=false2026.03 | 0.137 | — | — | 76.1 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=6.3%, Reasoning=false2026.03 | 0.129 | — | — | 75.7 | |
| Random DropBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.128 | — | — | 79.4 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.123 | — | — | 77.7 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=12.3%, Reasoning=false2026.03 | 0.113 | — | — | 77.9 | |
| Random DropBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.113 | — | — | 74.3 | |
| ToMeBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 0.111 | — | — | 79.1 | |
| ToMeBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.109 | — | — | 80.3 | |
| Random DropBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=25.0%, Reasoning=false2026.03 | 0.107 | — | — | 80.3 | |
| FrozenBiLMVision Encoder=ViT-L2024.02 | 0.071 | 0.1 | 0.044 | — | |
| Temp[BLIP]Vision Encoder=ViT-B2024.02 | 0.069 | 0.1 | 0.045 | — | |
| Temp[CLIP]Vision Encoder=ViT-B2024.02 | 0.061 | 0.083 | 0.037 | — | |
| Temp[Swin]Vision Encoder=SWT2024.02 | 0.049 | 0.066 | 0.023 | — | |
| VIOLETV2Vision Encoder=VSWT2024.02 | 0.031 | 0.043 | 0.013 | — | |
| VGTVision Encoder=RCNN2024.02 | 0.03 | 0.042 | 0.014 | — |