Forced-choice Visual Question Answering on Semantic-stress evaluation set
91.5Pair AccuracyInternVL3-2B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| InternVL3-2BPrecision=bfloat16, Max new tokens=4, Hardware=NVIDIA RTX A6000, Mode=Inference-only, Evaluation protocol=Deterministic forward-pass logits2026.06 | 91.5 | 84.9 | 100 | 100 | 100 | 0 | |
| Qwen3VL-2BPrecision=bfloat16, Max new tokens=4, Hardware=NVIDIA RTX A6000, Mode=Inference-only, Evaluation protocol=Deterministic forward-pass logits2026.06 | 91.4 | 83.8 | 0 | 0.5 | 49.9 | 17.3 | |
| Qwen3VL-8BPrecision=bfloat16, Max new tokens=4, Hardware=NVIDIA RTX A6000, Mode=Inference-only, Evaluation protocol=Deterministic forward-pass logits2026.06 | 90.8 | 83.9 | 47.8 | 37 | 0 | 16.2 | |
| Qwen3VL-4BPrecision=bfloat16, Max new tokens=4, Hardware=NVIDIA RTX A6000, Mode=Inference-only, Evaluation protocol=Deterministic forward-pass logits2026.06 | 89.1 | 80.1 | 92.1 | 14.8 | 76.7 | 78 | |
| Gemma3-4BPrecision=bfloat16, Max new tokens=4, Hardware=NVIDIA RTX A6000, Mode=Inference-only, Evaluation protocol=Deterministic forward-pass logits2026.06 | 88.7 | 78.7 | 9.2 | 0 | 33.2 | 100 |