Visual Question Answering on SLAKE, VQA-RAD, and PathVQA Error-specific Subsets
69.4MMDPO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DPOmodality=text-only, strategy=targeted preference construction2026.01 | 69.4 | 35.9 | 45.5 | 20.8 | |
| mDPOmodality=joint text-image, strategy=targeted preference construction2026.01 | 69.1 | 35.9 | 45.5 | 20.9 | |
| CoPOmodality=image-only, strategy=targeted preference construction2026.01 | 69 | 36 | 45.6 | 20.8 | |
| Text-Hallu2026.01 | 63.6 | 30.4 | 42.2 | 13.2 | |
| Text-Hallu + NLLloss=Negative Log-Likelihood2026.01 | 60.1 | 31.2 | 43 | 14.9 | |
| mDPO2026.01 | 59.5 | 32.1 | 42.7 | 14.7 | |
| MMedPO2026.01 | 56.7 | 32.9 | 41.5 | 12.3 | |
| SFT2026.01 | 56.2 | 30.2 | 43.2 | 14.2 | |
| Text-Noise + NLLloss=Negative Log-Likelihood2026.01 | 56 | 32.3 | 42.7 | 13.7 | |
| Image-ROI2026.01 | 55.5 | 30.3 | 42.6 | 11.7 | |
| Text-Noise2026.01 | 54.2 | 29.8 | 41.6 | 14.8 | |
| IRPO2026.01 | 54.1 | 31.8 | 40.8 | 13.9 | |
| Image-Noise2026.01 | 51.8 | 32.2 | 43.4 | 14.9 | |
| Base model2026.01 | 49.8 | 26.6 | 40.3 | 10 |