Visual Question Answering on A-OKVQA (test)
90.56AccuracyTTH (Ours)
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TTH (Ours)Target Model=Gemini 2.5 Flash Lite2026.05 | 90.56 | — | — | — | — | — | — | — | — | — | — | 61.73 | |
| External judgeTarget Model=GPT-5 Nano, Judge=GPT-5-Mini2026.05 | 89.43 | — | — | — | — | — | — | — | — | — | — | 44.69 | |
| V3Fusion-RectifyModel ID=1232026.03 | 89.17 | — | — | — | — | — | — | — | — | — | — | — | |
| V3Fusion-MLPModel ID=1232026.03 | 88.31 | — | — | — | — | — | — | — | — | — | — | — | |
| TTH (Ours)Target Model=GPT-5 Nano2026.05 | 88.12 | — | — | — | — | — | — | — | — | — | — | 45.81 | |
| V3Fusion-LEDModel ID=1232026.03 | 87.86 | — | — | — | — | — | — | — | — | — | — | — | |
| Self-refineTarget Model=Gemini 2.5 Flash Lite2026.05 | 87.69 | — | — | — | — | — | — | — | — | — | — | 35.71 | |
| Qwen2.5-VL-7b-InstructModel ID=52026.03 | 87.24 | — | — | — | — | — | — | — | — | — | — | — | |
| External judgeTarget Model=Gemini 2.5 Flash Lite, Judge=Gemini 2.5 Flash2026.05 | 86.81 | — | — | — | — | — | — | — | — | — | — | 43.37 | |
| MMBoundary w/o RCSAblation Stage=Reinforcement Learning Stage2025.05 | 85.7 | 34.3 | 36.8 | 64.8 | — | — | — | — | — | — | — | — | |
| Intern-VL2-8bModel ID=62026.03 | 85.32 | — | — | — | — | — | — | — | — | — | — | — | |
| Self-refineTarget Model=GPT-5 Nano2026.05 | 85.24 | — | — | — | — | — | — | — | — | — | — | 17.88 | |
| BaseTarget Model=GPT-5 Nano2026.05 | 84.37 | — | — | — | — | — | — | — | — | — | — | — | |
| TTH (Ours)Target Model=Claude 4.5 Haiku2026.05 | 84.27 | — | — | — | — | — | — | — | — | — | — | 49.4 | |
| CoTTarget Model=Gemini 2.5 Flash Lite2026.05 | 84.1 | — | — | — | — | — | — | — | — | — | — | 21.43 | |
| MMBoundary2025.05 | 83.5 | 31.6 | 30.4 | 66.1 | — | — | — | — | — | — | — | — | |
| LlaVA-v1.6-Vicuna-13bModel ID=12026.03 | 83.4 | — | — | — | — | — | — | — | — | — | — | — | |
| BaseTarget Model=Gemini 2.5 Flash Lite2026.05 | 82.88 | — | — | — | — | — | — | — | — | — | — | — | |
| CoTTarget Model=GPT-5 Nano2026.05 | 82.79 | — | — | — | — | — | — | — | — | — | — | 18.44 | |
| LlaVA-v1.6-Vicuna-7bModel ID=22026.03 | 82.62 | — | — | — | — | — | — | — | — | — | — | — | |
| MMBoundary w/o UMTEAblation Stage=Warm-Up Stage2025.05 | 82.4 | 33.7 | 33.2 | 65.3 | — | — | — | — | — | — | — | — | |
| MMBoundary w/o RECAblation Stage=Reinforcement Learning Stage2025.05 | 81.9 | 33.2 | 35.7 | 63.5 | — | — | — | — | — | — | — | — | |
| MMBoundary w/o ULNLPAblation Stage=Warm-Up Stage2025.05 | 81.5 | 32.7 | 34.3 | 64.2 | — | — | — | — | — | — | — | — | |
| MMBoundary w/o UTSARAblation Stage=Warm-Up Stage2025.05 | 81.3 | 32.4 | 35.8 | 62.7 | — | — | — | — | — | — | — | — | |
| External judgeTarget Model=Claude 4.5 Haiku, Judge=Claude 4.5 Sonnet2026.05 | 80.87 | — | — | — | — | — | — | — | — | — | — | 44.63 | |
| MMBoundary w/o UCLIPSAblation Stage=Warm-Up Stage2025.05 | 80.6 | 33.7 | 35.4 | 63.1 | — | — | — | — | — | — | — | — | |
| PaLI-X-VPDScale=55B, Type=specialist2023.12 | 80.4 | — | — | — | — | — | — | — | — | — | 68.2 | — | |
| CoTTarget Model=Claude 4.5 Haiku2026.05 | 80.26 | — | — | — | — | — | — | — | — | — | — | 26.1 | |
| MMBoundary w/o RKAAblation Stage=Reinforcement Learning Stage2025.05 | 80.2 | 32.5 | 34.7 | 62.9 | — | — | — | — | — | — | — | — | |
| DeepSeek-VL2-SmallModel ID=42026.03 | 80.08 | — | — | — | — | — | — | — | — | — | — | — | |
| MMBoundary w/o S-SMappingAblation Stage=Warm-Up Stage2025.05 | 79.3 | 34 | 36.2 | 63.4 | — | — | — | — | — | — | — | — | |
| RCE2025.05 | 78.8 | 36.1 | 39.4 | 62 | — | — | — | — | — | — | — | — | |
| Conf-CSR2025.05 | 78.5 | 40.8 | 43.7 | 61.8 | — | — | — | — | — | — | — | — | |
| BaseTarget Model=Claude 4.5 Haiku2026.05 | 78.25 | — | — | — | — | — | — | — | — | — | — | — | |
| Self-refineTarget Model=Claude 4.5 Haiku2026.05 | 77.9 | — | — | — | — | — | — | — | — | — | — | 30.52 | |
| MMBoundary w UMaxAblation Stage=Warm-Up Stage2025.05 | 77.4 | 36.2 | 38.6 | 58.3 | — | — | — | — | — | — | — | — | |
| MMBoundary w/o RLAblation Stage=Reinforcement Learning Stage2025.05 | 76.8 | 39.2 | 42.7 | 58.1 | — | — | — | — | — | — | — | — | |
| DeepSeek-VL2-TinyModel ID=32026.03 | 76.68 | — | — | — | — | — | — | — | — | — | — | — | |
| PaLI-3-VPDScale=5B, Type=specialist2023.12 | 76.5 | — | — | — | — | — | — | — | — | — | 63.6 | — | |
| DRL2025.05 | 74.6 | 39.5 | 45.3 | 61.4 | — | — | — | — | — | — | — | — | |
| SaySelf2025.05 | 73.4 | 34.5 | 38.4 | 60.3 | — | — | — | — | — | — | — | — | |
| InstructBLIP (Vicuna-7B)Backbone=Vicuna-7B2023.12 | 73.4 | — | — | — | — | — | — | — | — | — | 62.1 | — | |
| SC2025.05 | 70.1 | 43.5 | 49.2 | 54.8 | — | — | — | — | — | — | — | — | |
| Multisample2025.05 | 68.3 | 41.3 | 43 | 54.3 | — | — | — | — | — | — | — | — | |
| DPS2025.05 | 67.5 | 51.1 | 55.4 | 53.5 | — | — | — | — | — | — | — | — | |
| DPV2025.05 | 65 | 56.3 | 58.2 | 51.2 | — | — | — | — | — | — | — | — | |
| InstructBLIPLLM Backbone=Vicuna-7B, Evaluation Protocol=finetuning2023.05 | 62.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Previous SOTAEvaluation Protocol=finetuning2023.05 | 61.6 | — | — | — | — | — | — | — | — | — | — | — | |
| BLIP-2LLM Backbone=Vicuna-7B, Evaluation Protocol=finetuning2023.05 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | |
| FlamingoStatus=Unified SOTA2022.06 | 57.8 | — | — | — | — | — | — | — | — | — | — | — | |
| InstructBLIPLLM Backbone=FlanT5-XXL, Evaluation Protocol=finetuning2023.05 | 54.8 | — | — | — | — | — | — | — | — | — | — | — | |
| KATStatus=Fine-tuned SOTA2022.06 | 54.4 | — | — | — | — | — | — | — | — | — | — | — | |
| UNIFIED-IO XLModel Size=XL2022.06 | 54 | — | — | — | — | — | — | — | — | — | — | — | |
| BLIP-2LLM Backbone=FlanT5-XXL, Evaluation Protocol=finetuning2023.05 | 53.7 | — | — | — | — | — | — | — | — | — | — | — | |
| UNIFIED-IO LARGEModel Size=LARGE2022.06 | 42.7 | — | — | — | — | — | — | — | — | — | — | — | |
| GPV2Knowledge Sources=Web Search (Web10k) + COCO P.T., Approx. Params=220M2022.10 | 40.7 | — | — | — | — | — | — | — | — | — | — | — | |
| GPV-2Evaluation Protocol=Fine-Tuned2022.12 | 40.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=175B2022.12 | 40.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLM 175BEnd-to-End Training=false, Shot Number=02022.12 | 40.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=66B2022.12 | 38.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLM 66BEnd-to-End Training=false, Shot Number=02022.12 | 38.2 | — | — | — | — | — | — | — | — | — | — | — | |
| VLC-BERTKnowledge Sources=VQA P.T. + COMET, Approx. Params=118M2022.10 | 38.05 | — | — | — | — | — | — | — | — | — | — | — | |
| UNIFIED-IO BASEModel Size=BASE2022.06 | 37.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=30B2022.12 | 36 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLM 30BEnd-to-End Training=false, Shot Number=02022.12 | 36 | — | — | — | — | — | — | — | — | — | — | — | |
| Dual-EncoderModule Configuration=MuRes(V+L)2025.12 | 35.02 | — | — | — | — | — | — | — | — | — | — | — | |
| S³CEvaluation Protocol=unfiltered, Semi-supervised Learning=true2023.09 | 33.5 | — | — | — | 22.5 | 18.5 | 48.4 | 18.1 | 74.4 | 54.7 | — | — | |
| Dual-EncoderModule Configuration=With Residual2025.12 | 33.45 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=13B2022.12 | 33 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLM 13BEnd-to-End Training=false, Shot Number=02022.12 | 33 | — | — | — | — | — | — | — | — | — | — | — | |
| Dual-EncoderModule Configuration=MuRes(V)2025.12 | 32.89 | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPModule Configuration=With Residual2025.12 | 32.78 | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPModule Configuration=MuRes(V)2025.12 | 32.78 | — | — | — | — | — | — | — | — | — | — | — | |
| Dual-EncoderModule Configuration=Without Residual2025.12 | 32.64 | — | — | — | — | — | — | — | — | — | — | — | |
| VisualBERTModule Configuration=MuRes(V+L)2025.12 | 32.62 | — | — | — | — | — | — | — | — | — | — | — | |
| ViLTModule Configuration=MuRes(V+L)2025.12 | 32.53 | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPModule Configuration=MuRes(V+L)2025.12 | 32.47 | — | — | — | — | — | — | — | — | — | — | — | |
| VisualBERTModule Configuration=With Residual2025.12 | 32.47 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLMEvaluation Protocol=Zero-Shot, LLM Parameter Scale=6.7B2022.12 | 32.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Img2LLM 6.7BEnd-to-End Training=false, Shot Number=02022.12 | 32.2 | — | — | — | — | — | — | — | — | — | — | — | |
| ViLTModule Configuration=MuRes(V)2025.12 | 32.19 | — | — | — | — | — | — | — | — | — | — | — | |
| Dual-EncoderModule Configuration=MuRes(L)2025.12 | 31.72 | — | — | — | — | — | — | — | — | — | — | — | |
| ViLTModule Configuration=Without Residual2025.12 | 31.61 | — | — | — | — | — | — | — | — | — | — | — | |
| ViLTModule Configuration=MuRes(L)2025.12 | 31.48 | — | — | — | — | — | — | — | — | — | — | — | |
| ViLTModule Configuration=With Residual2025.12 | 31.21 | — | — | — | — | — | — | — | — | — | — | — | |
| VisualBERTModule Configuration=MuRes(L)2025.12 | 31.15 | — | — | — | — | — | — | — | — | — | — | — | |
| UNIFIED-IO SMALLModel Size=SMALL2022.06 | 31 | — | — | — | — | — | — | — | — | — | — | — | |
| VisualBERTModule Configuration=MuRes(V)2025.12 | 30.72 | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPModule Configuration=MuRes(L)2025.12 | 30.42 | — | — | — | — | — | — | — | — | — | — | — | |
| VisualBERTModule Configuration=Without Residual2025.12 | 29.88 | — | — | — | — | — | — | — | — | — | — | — | |
| S³C*Evaluation Protocol=unfiltered, Semi-supervised Learning=false2023.09 | 29.6 | — | — | — | 21.8 | 17.9 | 47.3 | 17.3 | 70.6 | 49.4 | — | — | |
| CLIPModule Configuration=Without Residual2025.12 | 29.41 | — | — | — | — | — | — | — | — | — | — | — | |
| NLX-GPTEvaluation Protocol=unfiltered2023.09 | 28.7 | — | — | — | 20.1 | 17 | 46.3 | 15.8 | 65.4 | 46.9 | — | — | |
| KRISPKnowledge Sources=Wikipedia + ConceptNet, Approx. Params=116M2022.10 | 27.1 | — | — | — | — | — | — | — | — | — | — | — | |
| KRISPEvaluation Protocol=Fine-Tuned2022.12 | 27.1 | — | — | — | — | — | — | — | — | — | — | — | |
| ViLBERTEvaluation Protocol=Fine-Tuned2022.12 | 25.9 | — | — | — | — | — | — | — | — | — | — | — | |
| LXMERTEvaluation Protocol=Fine-Tuned2022.12 | 25.9 | — | — | — | — | — | — | — | — | — | — | — | |
| LXMERT2022.10 | 25.89 | — | — | — | — | — | — | — | — | — | — | — | |
| VILBERTApprox. Params=116M2022.10 | 25.85 | — | — | — | — | — | — | — | — | — | — | — | |
| e-UGEvaluation Protocol=unfiltered2023.09 | 25.6 | — | — | — | 15.1 | 18.1 | 42.4 | 14.9 | 51.5 | 44.1 | — | — |