Visual Question Answering on GQA (testdev)
61.03Overall AccuracyHyperVis
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| HyperVisgeometry=hyperbolic, evaluation_mode=LoRA-only2026.06 | 61.03 | 46.52 | 81.88 | 80.96 | 76.1 | 64.52 | |
| LLaVA-1.5-7B + Eucl. visual relationsgeometry=Euclidean, variant=flat, evaluation_mode=LoRA-only2026.06 | 60.81 | 47.16 | 79.97 | 82.02 | 74.38 | 63.16 | |
| LLaVA-1.5-7Bstatus=baseline, evaluation_mode=LoRA-only2026.06 | 60.38 | 46.2 | 80.55 | 80.96 | 75.32 | 61.97 | |
| LLaVA-1.5-7B + textual SGG tripletsinput_extension=textual SGG triplets, evaluation_mode=LoRA-only2026.06 | 58.86 | — | — | — | — | — | |
| LLaVA-1.5-7B + LoRA onlyrelational_loss=none, evaluation_mode=LoRA-only2026.06 | 57.21 | 42.72 | 80.86 | 80.51 | 67.67 | 57.56 |