PET on AI2-THOR in-domain
97.8AccuracyBagel + Mixed Training
Evaluation Results
| Method | Links | |
|---|---|---|
| Bagel + Mixed TrainingModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=IPT + Label-only mixture2026.06 | 97.8 | |
| Bagel (label-only)Model Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Answer supervision only2026.06 | 97.5 | |
| Bagel + IPTModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Imaginative Perception Token2026.06 | 96.8 | |
| Bagel + Text CoTModel Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Fine-tuned, Reasoning Strategy=Textual chain-of-thought2026.06 | 83.1 | |
| GPT-5Model Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 79.8 | |
| Gemini 3 FlashModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 55 | |
| Qwen3-VL-8BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 52 | |
| Janus-Pro-7BModel Category=Unified Models, Evaluation Protocol=Zero-shot2026.06 | 51.8 | |
| InternVL3.5-8BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 51.5 | |
| Gemini 2.5 FlashModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 51 | |
| Qwen2.5-VL-7BModel Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 50.7 | |
| GPT-5.2Model Category=VQA Models, Evaluation Protocol=Zero-shot2026.06 | 45.5 | |
| Bagel (base)Model Category=Ours (fine-tuned BAGEL), Evaluation Protocol=Zero-shot2026.06 | 40.3 | |
| Chameleon 7BModel Category=Unified Models, Evaluation Protocol=Zero-shot2026.06 | 34.3 |