PT on Different-environment out-of-domain (Accuracy (%))
83.2Accuracy (OOD PT)Gemini 3 Flash
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 3 FlashModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 83.2 | |
| GPT-5Model Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 80.9 | |
| Gemini 2.5 FlashModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 71.4 | |
| Qwen3-VL-8BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 64.1 | |
| GPT-5.2Model Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 63 | |
| Bagel + Mixed TrainingModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=IPT + Label-only mixture2026.06 | 58.6 | |
| Bagel + IPTModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Imaginative Perception Token2026.06 | 57.5 | |
| Bagel (label-only)Model Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Answer supervision only2026.06 | 54.7 | |
| Bagel + Text CoTModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Textual chain-of-thought2026.06 | 52.2 | |
| InternVL3.5-8BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 47.4 | |
| Qwen2.5-VL-7BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 44.8 | |
| Bagel (base)Model Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Zero-shot2026.06 | 42.7 | |
| Janus-Pro-7BModel Category=Unified Models, Evaluation Protocol=Zero-shot2026.06 | 35.3 | |
| Chameleon 7BModel Category=Unified Models, Evaluation Protocol=Zero-shot2026.06 | 24.5 |