Visual Question Answering on M3CoT (test)
78.67Language Science Scoreperception-verified self-training
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| perception-verified self-trainingBackbone=Qwen-2-VL-Instruct, Evaluation Protocol=Self-Training2026.06 | 78.67 | 82.42 | 49.04 | 68.45 | |
| VQA FinetuneBackbone=Qwen-2-VL-Instruct, Evaluation Protocol=Direct SFT2026.06 | 76.78 | 80.22 | 45.22 | 65.39 | |
| R3VBackbone=Qwen-2-VL-Instruct, Evaluation Protocol=Self-Training2026.06 | 76.78 | 78.68 | 44.9 | 63.86 | |
| STaRBackbone=Qwen-2-VL-Instruct, Evaluation Protocol=Self-Training2026.06 | 75.83 | 78.02 | 46.5 | 66.67 | |
| Direct VQABackbone=Qwen-2-VL-Instruct, Evaluation Protocol=Zero-Shot2026.06 | 68.72 | 77.8 | 38.85 | 54.15 | |
| perception-verified self-trainingBackbone=LLaVA-v1.5-7B, Learning Strategy=Self-Training2026.06 | 64.93 | 72.74 | 46.56 | 58.24 | |
| [CAP-REAS-CONCL]Backbone=Qwen-2-VL-Instruct, Evaluation Protocol=Zero-Shot2026.06 | 63.03 | 54.72 | 37.58 | 47.13 | |
| R3VBackbone=LLaVA-v1.5-7B, Learning Strategy=Self-Training2026.06 | 50.24 | 65.49 | 42.2 | 51.34 | |
| STaRBackbone=LLaVA-v1.5-7B, Learning Strategy=Self-Training2026.06 | 48.82 | 64.98 | 41.88 | 53.9 | |
| VQA FinetuneBackbone=LLaVA-v1.5-7B, Learning Strategy=Direct SFT2026.06 | 46.45 | 60.22 | 34.24 | 46.1 | |
| Direct VQABackbone=LLaVA-v1.5-7B, Learning Strategy=Zero-Shot2026.06 | 45.5 | 57.58 | 29.62 | 36.4 | |
| [CAP-REAS-CONCL]Backbone=LLaVA-v1.5-7B, Learning Strategy=Zero-Shot2026.06 | 40.28 | 53.4 | 27.55 | 33.2 | |
| CoTBackbone=LLaVA-v1.5-7B, Learning Strategy=Zero-Shot2026.06 | 38.86 | 59.34 | 25.48 | 33.59 | |
| Desp-CoTBackbone=LLaVA-v1.5-7B, Learning Strategy=Zero-Shot2026.06 | 34.12 | 54.73 | 25.32 | 32.18 | |
| CCoTBackbone=LLaVA-v1.5-7B, Learning Strategy=Zero-Shot2026.06 | 26.54 | 54.08 | 28.66 | 35.5 |