Medical Difference Visual Question Answering on MIMIC-Diff-VQA (test)
0.594BLEU-4Location Aware Pretraining
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Location Aware Pretrainingdecoder=GPT-22026.03 | 0.594 | 0.425 | 0.747 | 2.997 | 0.972 | |
| RG-AGdecoder=GPT-22026.03 | 0.551 | 0.384 | 0.668 | 2.198 | 0.965 | |
| ReAldecoder=GPT-22026.03 | 0.53 | 0.395 | 0.736 | 2.409 | 0.968 | |
| PLURALdecoder=GPT-22026.03 | 0.52 | 0.381 | 0.653 | 1.832 | 0.963 | |
| BLIP-2decoder=GPT-22026.03 | 0.375 | 0.35 | 0.545 | 0.801 | 0.96 | |
| Regional Contrastive Pretrainingdecoder=GPT-22026.03 | 0.36 | 0.338 | 0.465 | 0.635 | 0.956 | |
| CapPadecoder=GPT-22026.03 | 0.35 | 0.327 | 0.529 | 0.675 | 0.948 | |
| Global Contrastive Pretrainingdecoder=GPT-22026.03 | 0.291 | 0.3 | 0.407 | 0.55 | 0.85 |