Video Understanding on HERBench Lite
66.7AccuracyVideoChat-A1
Evaluation Results
| Method | Links | |
|---|---|---|
| VideoChat-A1Model=GPT-4o, Frame Budget (K)=384, Evaluation Protocol=Stratified random 25% subset2026.03 | 66.7 | |
| HiMuModel=Qwen3-VL-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 43.22 | |
| BOLTModel=Qwen3-VL-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 42.2 | |
| Uniform SamplingModel=Qwen3-VL-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 41.7 | |
| HiMuModel=GPT-4o, Frame Budget (K)=16, Evaluation Protocol=Stratified random 25% subset2026.03 | 40.68 | |
| AKSModel=Qwen3-VL-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 40.25 | |
| T*Model=Qwen3-VL-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 39.1 | |
| HiMuModel=InternVL-3.5-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 38.32 | |
| UniformModel=InternVL-3.5-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 38.3 | |
| HiMuModel=Gemini-2.5-Flash, Frame Budget (K)=16, Evaluation Protocol=Stratified random 25% subset2026.03 | 37.68 | |
| UniformModel=GPT-4o, Frame Budget (K)=16, Evaluation Protocol=Stratified random 25% subset2026.03 | 37.47 | |
| UniformModel=Gemini-2.5-Flash, Frame Budget (K)=16, Evaluation Protocol=Stratified random 25% subset2026.03 | 37.27 | |
| HiMuModel=LLaVA-OV-1.5-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 35.87 | |
| UniformModel=LLaVA-OV-1.5-8B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 35.75 | |
| HiMuModel=Qwen2.5-VL-7B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 35.17 | |
| UniformModel=Qwen2.5-VL-7B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 34.05 | |
| HiMuModel=Gemma-3-12B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 31.47 | |
| UniformModel=Gemma-3-12B, Frame Budget (K)=16, Evaluation Protocol=Full benchmark2026.03 | 31.2 |