Visual Reasoning on NLVR2 (test)
85.15AccuracySimVLM_HUGE
Evaluation Results
| Method | Links | |
|---|---|---|
| SimVLM_HUGEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 85.15 | |
| VinVL_LARGEPre-training scale=<10M2021.11 | 83.98 | |
| X-VLMclip2022.10 | 83.48 | |
| MADTPReduce Ratio=0.3, GFLOPS=92.60, Backbone=BLIP2024.03 | 83.2 | |
| ALBEFPre-training volume=14M, Pre-training scale=>10M2021.11 | 83.14 | |
| UncompressedReduce Ratio=/, GFLOPS=132.54, Backbone=BLIP2024.03 | 83.08 | |
| METERBackbone=CLIP-ViT-BASE, Pre-training scale=<10M2021.11 | 83.05 | |
| MADTPReduce Ratio=0.5, GFLOPS=66.16, Backbone=BLIP2024.03 | 82.85 | |
| METERBackbone=Swin-BASE, Pre-training scale=<10M2021.11 | 82.47 | |
| MADTPReduce Ratio=0.6, GFLOPS=52.92, Backbone=BLIP2024.03 | 82.42 | |
| X-VLMclipperformance_target=98%2022.10 | 81.81 | |
| SimVLM_BASEPre-training volume=1.8B, Pre-training scale=>10M2021.11 | 81.77 | |
| EfficientVLM2022.10 | 81.72 | |
| VILLA_LARGEPre-training scale=<10M2021.11 | 81.47 | |
| MADTPReduce Ratio=0.7, GFLOPS=39.69, Backbone=BLIP2024.03 | 81.23 | |
| EVPGFramework=VisProg, Supervision Type=Supervised, Backbone=BLIP_lavis2025.12 | 81.2 | |
| UPopReduce Ratio=0.3, GFLOPS=89.36, Backbone=BLIP2024.03 | 81.13 | |
| ALBEFPre-training volume=4M, Pre-training scale=<10M2021.11 | 80.5 | |
| EVPGFramework=VisProg, Supervision Type=Supervised, Backbone=InternVL1.52025.12 | 80.5 | |
| STPReduce Ratio=0.3, GFLOPS=94.08, Backbone=BLIP2024.03 | 80.01 | |
| UNITER_LARGEPre-training scale=<10M2021.11 | 79.98 | |
| CoMPGFLOPs=26.33±0.382026.04 | 79.67 | |
| X-VLMsmall2022.10 | 79.26 | |
| MADTPReduce Ratio=0.8, GFLOPS=26.46, Backbone=BLIP2024.03 | 79.22 | |
| OSCARB2022.10 | 78.36 | |
| InternVL1.5Framework=VisProg, Supervision Type=Zero-shot (baseline)2025.12 | 78.2 | |
| Visual ParsingPre-training scale=<10M2021.11 | 78.05 | |
| MADTPGFLOPs=26.77±0.232026.04 | 77.64 | |
| STPReduce Ratio=0.5, GFLOPS=68.31, Backbone=BLIP2024.03 | 77.61 | |
| UPopReduce Ratio=0.5, GFLOPS=65.29, Backbone=BLIP2024.03 | 77.61 | |
| OSCARBperformance_target=98%2022.10 | 76.79 | |
| VILT-NLVRFinetuned=true, Runs=12022.11 | 76.3 | |
| ViLTPre-training scale=<10M2021.11 | 76.13 | |
| ViLT2022.10 | 76.1 | |
| SDVPFramework=VisProg, Supervision Type=Supervised, Backbone=BLIP_hg2025.12 | 75.2 | |
| DistilDualEnc2022.10 | 74.3 | |
| MiniVLM2022.10 | 73.93 | |
| BLIP_lavisFramework=VisProg, Supervision Type=Zero-shot (baseline)2025.12 | 73.6 | |
| UPopReduce Ratio=0.6, GFLOPS=50.35, Backbone=BLIP2024.03 | 73.55 | |
| PixelBERTPre-training scale=<10M2021.11 | 72.2 | |
| UPopReduce Ratio=0.7, GFLOPS=39.93, Backbone=BLIP2024.03 | 68.76 | |
| BLIP_hgFramework=VisProg, Supervision Type=Zero-shot (baseline)2025.12 | 68.3 | |
| VISPROGPrompting strategy=voting, Finetuned=false, Runs=5, Context examples per run=162022.11 | 62.4 | |
| VISPROGPrompting strategy=curated, Finetuned=false, Runs=1, Context examples per run=122022.11 | 61.8 | |
| VISPROGPrompting strategy=random, Finetuned=false, Runs=1, Context examples per run=162022.11 | 61.3 | |
| UPopReduce Ratio=0.8, GFLOPS=19.08, Backbone=BLIP2024.03 | 57.79 |