Multimodal Understanding on LLaVA-Bench In-the-Wild
5.833AccuracyILVAD
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ILVADBackbone=LLaVA-1.5-7B, τ=5, α=5, Decoding=Greedy2026.05 | 5.833 | 6.133 | 6.867 | |
| ILVADBackbone=Qwen2-VL-7B, τ=5, α=5, Decoding=Greedy2026.05 | 5.558 | 5.94 | 6.197 | |
| ILVADBackbone=LLaVA-NeXT-7B, τ=5, α=3, Decoding=Greedy2026.05 | 5.467 | 6.633 | 7.633 | |
| LLaVA-1.5-7BDecoding=Greedy2026.05 | 5.45 | 5.833 | 7.017 | |
| LLaVA-NeXT-7BDecoding=Greedy2026.05 | 5.4 | 6.417 | 7.767 | |
| Qwen2-VL-7BDecoding=Greedy2026.05 | 5.233 | 5.992 | 6.11 |