Temporal Grounding on ActivityNet
74.1Recall@0.3VideoAuto-R1
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| VideoAuto-R1Backbone=Qwen3-VL-8B2026.01 | 74.1 | 54.3 | 32.4 | 51.9 | — | |
| VITAL2026.01 | 70.9 | 50.8 | 31.6 | 49.8 | — | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B2026.01 | 69.2 | 48.5 | 27.3 | 47.6 | — | |
| TimeMarker2026.01 | 67.4 | 50.7 | 33 | 49.5 | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=16.1%, Reasoning=true2026.03 | 65.8 | 44.9 | 23.8 | 44.7 | — | |
| Time-R12026.01 | 58.6 | 39 | 21.4 | 40.5 | — | |
| SER-7BModel Category=Grounded Reasoning Models, Protocol=Zero-shot2026.06 | 58.1 | 38.1 | 19 | 39.2 | — | |
| Temporal-RLT2026.01 | 56.9 | 38.4 | 20.2 | 39 | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=6.8%, Reasoning=true2026.03 | 53.4 | 34 | 16.4 | 35.7 | — | |
| VideoChat-R1.52026.01 | 52.4 | 32.3 | 16.8 | 35.3 | — | |
| FlashVidBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=31.3%, Reasoning=false2026.03 | 51.9 | 33.4 | 19 | 36.8 | — | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=100%, Reasoning=true2026.03 | 50.8 | 34.1 | 17.4 | 34.4 | — | |
| FlashVidBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=12.6%, Reasoning=false2026.03 | 49.9 | 31.5 | 17.4 | 35.2 | — | |
| Open-o3-VideoModel Category=Grounded Reasoning Models, Protocol=Zero-shot2026.06 | 49.5 | 30.8 | 15.9 | 34.4 | — | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=100%, Reasoning=true2026.03 | 49.4 | 34.3 | 18.5 | 33.5 | — | |
| HawkEyeModel Category=Temporal Grounding Video LLM, Protocol=Zero-shot2026.06 | 49.1 | 29.3 | 10.7 | 32.7 | — | |
| VideoMindModel Category=Temporal Grounding Video LLM, Protocol=Zero-shot2026.06 | 48.4 | 30.3 | 15.7 | 33.3 | — | |
| VanillaBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=100%, Reasoning=false2026.03 | 47.9 | 30.9 | 17.5 | 34.4 | — | |
| ToMeBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=10.0%, Reasoning=false2026.03 | 46.3 | 31 | 19.2 | 34.1 | — | |
| ToMeBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=25.0%, Reasoning=false2026.03 | 45.9 | 28.8 | 15.6 | 32.6 | — | |
| VanillaBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=100%, Reasoning=false2026.03 | 45.8 | 31.1 | 19.2 | 33.9 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 52026.05 | 45.35 | 28.01 | 16.38 | 32.59 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=52026.05 | 45.35 | 28.01 | 16.38 | 32.59 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 42026.05 | 44.76 | 27.67 | 16.24 | 32.12 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=42026.05 | 44.76 | 27.67 | 16.24 | 32.12 | — | |
| VanillaBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=100%, Reasoning=false2026.03 | 44.6 | 28.3 | 15.5 | 31.8 | — | |
| TimeChat2026.01 | 44 | 27.8 | 14.3 | 30.4 | — | |
| VTimeLLMModel Category=Temporal Grounding Video LLM, Protocol=Zero-shot2026.06 | 44 | 27.8 | 14.3 | 30.4 | — | |
| MomentorModel Category=Temporal Grounding Video LLM, Protocol=Zero-shot2026.06 | 42.9 | 23 | 12.4 | 29.3 | — | |
| ToMeBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 42.4 | 27.6 | 16.6 | 31.4 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 32026.05 | 41.82 | 26.15 | 15.5 | 30.29 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=32026.05 | 41.82 | 26.15 | 15.5 | 30.29 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 42026.05 | 41.77 | 26.23 | 14.66 | 29.87 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=42026.05 | 41.77 | 26.23 | 14.66 | 29.87 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 52026.05 | 41 | 25.83 | 14.39 | 29.42 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=52026.05 | 41 | 25.83 | 14.39 | 29.42 | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=16.1%, Reasoning=false2026.03 | 40.6 | 26.7 | 15.7 | 30 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 32026.05 | 40.17 | 25.2 | 14.15 | 28.99 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=32026.05 | 40.17 | 25.2 | 14.15 | 28.99 | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=16.2%, Reasoning=false2026.03 | 40 | 24.4 | 13 | 28.5 | — | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=12.3%, Reasoning=false2026.03 | 39.9 | 26.2 | 15.3 | 29.5 | — | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=12.3%, Reasoning=false2026.03 | 39.6 | 24.2 | 13.1 | 28.4 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 22026.05 | 38.49 | 23.94 | 13.42 | 27.79 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=22026.05 | 38.49 | 23.94 | 13.42 | 27.79 | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=6.8%, Reasoning=false2026.03 | 38.3 | 24.5 | 14.4 | 28.4 | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, reproduced=true2026.01 | 37.9 | 22.6 | 10.6 | 26.9 | — | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=6.3%, Reasoning=false2026.03 | 37.9 | 24.3 | 14.3 | 28.1 | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=6.8%, Reasoning=false2026.03 | 37.5 | 22.5 | 12.3 | 27.2 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 22026.05 | 37.48 | 23.16 | 13.66 | 27.3 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=22026.05 | 37.48 | 23.16 | 13.66 | 27.3 | — | |
| VisPlayBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 22026.05 | 37.28 | 23.11 | 12.94 | 26.85 | — | |
| V-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 32026.05 | 37.21 | 23.23 | 12.95 | 26.86 | — | |
| V-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 12026.05 | 37.16 | 23.29 | 13.11 | 26.89 | — | |
| FixedScaleBackbone=Qwen3-VL-8B, Frames=32, Retention Ratio R=6.3%, Reasoning=false2026.03 | 37 | 22.3 | 12 | 27 | — | |
| VisPlayBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 42026.05 | 36.94 | 22.99 | 12.93 | 26.74 | — | |
| VisPlayBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 52026.05 | 36.94 | 23.18 | 13.03 | 26.75 | — | |
| VisPlayBackbone=Qwen3-VL-8B-Instruct2026.05 | 36.94 | 22.99 | 12.93 | 26.74 | — | |
| V-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 22026.05 | 36.88 | 23.01 | 13.09 | 26.76 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 12026.05 | 36.87 | 23.08 | 12.99 | 26.75 | — | |
| Video-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=12026.05 | 36.87 | 23.08 | 12.99 | 26.75 | — | |
| VisPlayBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 32026.05 | 36.85 | 22.99 | 12.73 | 26.57 | — | |
| V-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 52026.05 | 36.83 | 23.25 | 13.12 | 26.83 | — | |
| V-ZeroBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 42026.05 | 36.79 | 23.05 | 13 | 26.69 | — | |
| V-ZeroBackbone=Qwen3-VL-8B-Instruct2026.05 | 36.79 | 23.05 | 13 | 26.69 | — | |
| VisPlayBackbone=Qwen3-VL-8B-Instruct, Iteration=Iter 12026.05 | 36.68 | 23.13 | 12.88 | 26.66 | — | |
| Base ModelBackbone=Qwen3-VL-8B-Instruct, Iteration=Base2026.05 | 36.59 | 22.63 | 12.85 | 26.56 | — | |
| Base ModelBackbone=Qwen3-VL-8B-Instruct2026.05 | 36.59 | 22.63 | 12.85 | 26.56 | — | |
| Random DropBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=25.0%, Reasoning=false2026.03 | 36.1 | 21.1 | 12.7 | 26.3 | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=6.8%, Reasoning=true2026.03 | 35.4 | 21.5 | 10 | 24.4 | — | |
| Random DropBackbone=Qwen3-VL-8B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 33.5 | 18.6 | 11.5 | 24.8 | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=16.1%, Reasoning=false2026.03 | 33.1 | 19.3 | 10.2 | 24.3 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 12026.05 | 32.97 | 20.22 | 12.04 | 24.23 | — | |
| Video-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=12026.05 | 32.97 | 20.22 | 12.04 | 24.23 | — | |
| V-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 52026.05 | 31.5 | 19.22 | 11.41 | 23.12 | — | |
| V-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 12026.05 | 31.24 | 19.17 | 11.4 | 23.07 | — | |
| VisPlayBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 12026.05 | 31.19 | 19.08 | 11.17 | 22.93 | — | |
| VisPlayBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 32026.05 | 31.05 | 19.14 | 11.19 | 22.83 | — | |
| VisPlayBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 52026.05 | 30.99 | 18.97 | 11.2 | 22.81 | — | |
| VisPlayBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 42026.05 | 30.98 | 19.12 | 11.31 | 22.89 | — | |
| VisPlayBackbone=Qwen3-VL-4B-Instruct2026.05 | 30.98 | 19.12 | 11.31 | 22.89 | — | |
| Base ModelBackbone=Qwen3-VL-4B-Instruct, Iteration=Base2026.05 | 30.94 | 19.04 | 11.29 | 22.86 | — | |
| Base ModelBackbone=Qwen3-VL-4B-Instruct2026.05 | 30.94 | 19.04 | 11.29 | 22.86 | — | |
| V-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 42026.05 | 30.91 | 18.92 | 11.31 | 22.81 | — | |
| V-ZeroBackbone=Qwen3-VL-4B-Instruct2026.05 | 30.91 | 18.92 | 11.31 | 22.81 | — | |
| V-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 22026.05 | 30.88 | 19.05 | 11.45 | 22.9 | — | |
| VisPlayBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 22026.05 | 30.84 | 18.82 | 11.12 | 22.76 | — | |
| V-ZeroBackbone=Qwen3-VL-4B-Instruct, Iteration=Iter 32026.05 | 30.55 | 18.85 | 11.28 | 22.63 | — | |
| VanillaBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=100%, Reasoning=false2026.03 | 30.4 | 18 | 8.9 | 22.6 | — | |
| ToMeBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=25.0%, Reasoning=false2026.03 | 27.2 | 14.4 | 6.4 | 19.1 | — | |
| Random DropBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=25.0%, Reasoning=false2026.03 | 26.7 | 13.9 | 6.3 | 18.8 | — | |
| Video-ChatGPTModel Category=General Video LLMs, Protocol=Zero-shot2026.06 | 26.4 | 13.6 | 6.1 | 18.9 | — | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=12.3%, Reasoning=false2026.03 | 25 | 13.8 | 5.9 | 18.3 | — | |
| Random DropBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 23.8 | 12 | 5.3 | 17 | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=6.8%, Reasoning=false2026.03 | 23.5 | 12.9 | 6.1 | 17.2 | — | |
| ToMeBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=10.0%, Reasoning=false2026.03 | 22.9 | 11.8 | 5.5 | 16.4 | — | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=128, Retention Ratio R=6.3%, Reasoning=false2026.03 | 22.8 | 12.8 | 5.7 | 17.1 | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=16.2%, Reasoning=false2026.03 | 19.8 | 10.8 | 5.2 | 15.3 | — | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=25.0%, Reasoning=false2026.03 | 18.6 | 9.4 | 4.3 | 14.1 | — | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=12.3%, Reasoning=false2026.03 | 17.5 | 8.9 | 4 | 13.3 | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Frames=32, Retention Ratio R=6.8%, Reasoning=false2026.03 | 16.3 | 8.5 | 3.9 | 12.5 | — |