PT on AI2-THOR in-domain
66.7AccuracyBagel + Mixed Training
Evaluation Results
| Method | Links | |
|---|---|---|
| Bagel + Mixed TrainingModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=IPT + Label-only mixture2026.06 | 66.7 | |
| Bagel (label-only)Model Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Answer supervision only2026.06 | 65.7 | |
| GPT-5Model Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 60.2 | |
| Bagel + Text CoTModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Textual chain-of-thought2026.06 | 49.7 | |
| Bagel + IPTModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Imaginative Perception Token2026.06 | 49 | |
| Gemini 3 FlashModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 42.3 | |
| Gemini 2.5 FlashModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 41.5 | |
| Qwen2.5-VL-7BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 37.3 | |
| Qwen3-VL-8BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 35.9 | |
| InternVL3.5-8BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 35.8 | |
| Janus-Pro-7BModel Category=Unified Models, Evaluation Protocol=Zero-shot2026.06 | 33.5 | |
| GPT-5.2Model Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 32.9 | |
| Bagel (base)Model Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Zero-shot2026.06 | 29.9 | |
| Chameleon 7BModel Category=Unified Models, Evaluation Protocol=Zero-shot2026.06 | 16.3 |