Visual Question Answering on OKVQA (VQA accuracy)
59.57VQA AccuracyTPC
Evaluation Results
| Method | Links | |
|---|---|---|
| TPCIntegration Strategy=Ours (core–residual alignment), Base LVLM=LLaVA-1.5-7B2026.07 | 59.57 | |
| CLIPPerturbation Budget=clean, Attack=APGD2025.02 | 59.04 | |
| Patch-Aligned TrainingIntegration Strategy=Projector / fine-grained alignment (LVLM-side training), Base LVLM=LLaVA-1.5-7B2026.07 | 58.3 | |
| Adv-W2SIntegration Strategy=Robust vision encoder replacement, Base LVLM=LLaVA-1.5-7B2026.07 | 56.8 | |
| Robust-LLaVA4_GPerturbation Budget=clean, Attack=APGD2025.02 | 54.88 | |
| FARE4Perturbation Budget=clean, Attack=APGD2025.02 | 54.64 | |
| FAREIntegration Strategy=Robust vision encoder replacement, Base LVLM=LLaVA-1.5-7B2026.07 | 54.1 | |
| ImageNet Std ViT-LPerturbation Budget=clean, Attack=APGD2025.02 | 53.64 | |
| LLaVA1.5 (Regular)Integration Strategy=Decoding / concept-level enhancement (frozen LVLM), Base LVLM=LLaVA-1.5-7B2026.07 | 53.4 | |
| Sim-CLIP4Perturbation Budget=clean, Attack=APGD2025.02 | 53.04 | |
| VL-SAEIntegration Strategy=Decoding / concept-level enhancement (frozen LVLM), Base LVLM=LLaVA-1.5-7B2026.07 | 53 | |
| PMG-AFTIntegration Strategy=Robust vision encoder replacement, Base LVLM=LLaVA-1.5-7B2026.07 | 52.6 | |
| TGA-ZSRIntegration Strategy=Robust vision encoder replacement, Base LVLM=LLaVA-1.5-7B2026.07 | 52 | |
| VCDIntegration Strategy=Decoding / concept-level enhancement (frozen LVLM), Base LVLM=LLaVA-1.5-7B2026.07 | 51.9 | |
| TeCoAIntegration Strategy=Robust vision encoder replacement, Base LVLM=LLaVA-1.5-7B2026.07 | 51.8 | |
| Robust-LLaVA4_GPerturbation Budget=2/255, Attack=APGD2025.02 | 49.68 | |
| ImageNet Robust ViT-BPerturbation Budget=clean, Attack=APGD2025.02 | 44.96 | |
| Robust-LLaVA4_GPerturbation Budget=4/255, Attack=APGD2025.02 | 43.76 | |
| ImageNet Robust ViT-BPerturbation Budget=2/255, Attack=APGD2025.02 | 39.36 | |
| Sim-CLIP4Perturbation Budget=2/255, Attack=APGD2025.02 | 36.2 | |
| Robust-LLaVA4_GPerturbation Budget=8/255, Attack=APGD2025.02 | 34.04 | |
| FARE4Perturbation Budget=2/255, Attack=APGD2025.02 | 32.32 | |
| ImageNet Robust ViT-BPerturbation Budget=4/255, Attack=APGD2025.02 | 31.92 | |
| Sim-CLIP4Perturbation Budget=4/255, Attack=APGD2025.02 | 30.88 | |
| FARE4Perturbation Budget=4/255, Attack=APGD2025.02 | 29.88 | |
| FARE4Perturbation Budget=8/255, Attack=APGD2025.02 | 24.52 | |
| ImageNet Robust ViT-BPerturbation Budget=8/255, Attack=APGD2025.02 | 23.48 | |
| Sim-CLIP4Perturbation Budget=8/255, Attack=APGD2025.02 | 22.96 | |
| CLIPPerturbation Budget=2/255, Attack=APGD2025.02 | 16.64 | |
| CLIPPerturbation Budget=4/255, Attack=APGD2025.02 | 13.88 | |
| CLIPPerturbation Budget=8/255, Attack=APGD2025.02 | 11.56 | |
| ImageNet Std ViT-LPerturbation Budget=2/255, Attack=APGD2025.02 | 11.16 | |
| ImageNet Std ViT-LPerturbation Budget=4/255, Attack=APGD2025.02 | 8.52 | |
| ImageNet Std ViT-LPerturbation Budget=8/255, Attack=APGD2025.02 | 6.36 |