Answerability on VizWiz
61.5AccuracyInternVL3.5-REGATE
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL3.5-REGATELLM=Qwen3-14B, Tokens=2.32B (↓ 41.41%), Zero-shot=true2025.07 | 61.5 | |
| VILA1.5LLM=Llama-2-13B, Tokens=–, Zero-shot=true2025.07 | 60.6 | |
| InternVL3.5LLM=Qwen3-14B, Tokens=3.96B, Zero-shot=true2025.07 | 60.6 | |
| LLaVA-1.6LLM=Vicuna-7B, Tokens=–, Zero-shot=true2025.07 | 57.6 | |
| LLaVA-OneVisionLLM=Qwen2-7B, Tokens=–, Zero-shot=true2025.07 | 53 | |
| LLaVA-1.5LLM=Vicuna-7B, Tokens=–, Zero-shot=true2025.07 | 50 | |
| VideoLLaMA2-REGATELLM=Qwen2-7B, Tokens=49.27M (↓ 41.22%), Zero-shot=true2025.07 | 48 | |
| VideoLLaMA2LLM=Qwen2-7B, Tokens=83.82M, Zero-shot=true2025.07 | 46.8 | |
| Qwen-VL-ChatLLM=Qwen-7B, Tokens=–, Zero-shot=true2025.07 | 38.9 | |
| InstructBLIPLLM=Vicuna-7B, Tokens=–, Zero-shot=true2025.07 | 34.5 | |
| VideoChat2-REGATELLM=Mistral-7B, Tokens=2.22B (↓ 43.51%), Zero-shot=true2025.07 | 32.5 | |
| VideoChat2LLM=Mistral-7B, Tokens=3.93B, Zero-shot=true2025.07 | 28.5 |