Visual Question Answering on (SLAKE, VQA-RAD, and PathVQA Pooled)
45.6AccuracyCoPO
Evaluation Results
| Method | Links | |
|---|---|---|
| CoPOmodality=image-only, strategy=targeted preference construction2026.01 | 45.6 | |
| DPOmodality=text-only, strategy=targeted preference construction2026.01 | 45.5 | |
| mDPOmodality=joint text-image, strategy=targeted preference construction2026.01 | 45.4 | |
| Text-Hallu + NLLloss=Negative Log-Likelihood2026.01 | 42 | |
| mDPO2026.01 | 41.9 | |
| Text-Hallu2026.01 | 41.3 | |
| SFT2026.01 | 41 | |
| Image-ROI2026.01 | 41 | |
| Text-Noise + NLLloss=Negative Log-Likelihood2026.01 | 40.6 | |
| MMedPO2026.01 | 40.1 | |
| Image-Noise2026.01 | 40 | |
| Text-Noise2026.01 | 39 | |
| IRPO2026.01 | 39 | |
| Base model2026.01 | 38.7 |