Video Question Answering on LongVideoBench (LVB) 10-hour variant 1.0 (test)
70.1AccuracyVideo-RLM
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Video-RLMbackbone=Gemini2026.03 | 70.1 | -1.9 | 307 | |
| LLM over Captionscaptions=GPT-4o, reasoning_model=Qwen3.52026.03 | 62.1 | -0.3 | 207 | |
| Qwen3.5sampling=uniform2026.03 | 49.2 | -12.3 | 212 | |
| Video-RLMbackbone=Qwen2026.03 | 47.7 | -4.8 | 146 |