Multimodal Evaluation on SEED-Bench
77.3AccuracyD2Dloc
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| D2DlocReasoning Strategy=D2D, Variant=loc, Backbone=Qwen2.5-VL-7B2025.07 | 77.3 | — | — | — | — | |
| Jigsaw2025.12 | 77.01 | — | — | — | — | |
| D2DjusReasoning Strategy=D2D, Variant=jus, Backbone=Qwen2.5-VL-7B2025.07 | 77 | — | — | — | — | |
| ViCrit2025.12 | 76.64 | — | — | — | — | |
| D2IlocReasoning Strategy=D2I, Variant=loc, Backbone=Qwen2.5-VL-7B2025.07 | 76.6 | — | — | — | — | |
| Qwen2.5-VL-7B w/ GRPOReasoning Mode=Deliberate, Backbone=Qwen2.5-VL-7B2025.07 | 76.5 | — | — | — | — | |
| D2DparReasoning Strategy=D2D, Variant=par, Backbone=Qwen2.5-VL-7B2025.07 | 76.5 | — | — | — | — | |
| D2IjusReasoning Strategy=D2I, Variant=jus, Backbone=Qwen2.5-VL-7B2025.07 | 76.4 | — | — | — | — | |
| Mix + CL + CARE2025.12 | 76.38 | — | — | — | — | |
| GRPO-CARE2025.12 | 76.36 | — | — | — | — | |
| Qwen2.5-VL-7B w/ GRPO†Reasoning Mode=Intuitive, Backbone=Qwen2.5-VL-7B2025.07 | 76.3 | — | — | — | — | |
| InternVL2-8BModel=InternVL2-8B, Pruning Method=None2025.09 | 76.1 | — | — | — | — | |
| D2IparReasoning Strategy=D2I, Variant=par, Backbone=Qwen2.5-VL-7B2025.07 | 76.1 | — | — | — | — | |
| Rotation + CL + CARE2025.12 | 76.05 | — | — | — | — | |
| Qwen2.5-VL-7B*Backbone=Qwen2.5-VL-7B2025.07 | 75.6 | — | — | — | — | |
| Qwen 2.5 VL2025.12 | 75.48 | — | — | — | — | |
| VisualSphinx2025.12 | 75.47 | — | — | — | — | |
| Vision-Zero2025.12 | 75.45 | — | — | — | — | |
| Jigsaw + CARE2025.12 | 75.34 | — | — | — | — | |
| Jigsaw + CL + CARE2025.12 | 75.17 | — | — | — | — | |
| PTPModel=InternVL2-8B, Pruning Method=PTP, Pruning Ratio (r)=0.5, Balancing Weight (alpha)=0.52025.09 | 75.1 | — | — | — | — | |
| Qwen2-VL-7B2025.07 | 75.1 | — | — | — | — | |
| Visual Jigsaw2025.12 | 74.8 | — | — | — | — | |
| Jigsaw + CL2025.12 | 74.64 | — | — | — | — | |
| PatchFit + CL + CARE2025.12 | 73.97 | — | — | — | — | |
| GPT-4o2025.07 | 72 | — | — | — | — | |
| Full FinetuneModel=Llama-3-8B, Sampling Ratio=100%2025.03 | 71.8 | — | — | — | — | |
| InternVL2-2BModel=InternVL2-2B, Pruning Method=None2025.09 | 70.9 | — | — | — | — | |
| PTPModel=InternVL2-2B, Pruning Method=PTP, Pruning Ratio (r)=0.5, Balancing Weight (alpha)=0.52025.09 | 70.7 | — | — | — | — | |
| LongVILA-7B (S3)LLM=Qwen2-7B, Resolution=dynamic2024.08 | 70.6 | — | — | — | — | |
| OursModel=Kimi-VL A3B-Thinking2025.10 | 69.74 | — | — | — | 53.62 | |
| OursModel=R1-Onevision 7B2025.10 | 69.52 | — | — | — | 53.01 | |
| AGLAModel=Kimi-VL A3B-Thinking2025.10 | 69.27 | — | — | — | 51.48 | |
| VCDModel=R1-Onevision 7B2025.10 | 68.76 | — | — | — | 52.31 | |
| PreSelModel=Llama-3-8B, Sampling Ratio=15%2025.03 | 68.5 | — | — | — | — | |
| VanillaModel=R1-Onevision 7B2025.10 | 68.48 | — | — | — | 52.06 | |
| RandomModel=Llama-3-8B, Sampling Ratio=15%2025.03 | 68.3 | — | — | — | — | |
| AGLAModel=R1-Onevision 7B2025.10 | 68.11 | — | — | — | 51.65 | |
| Full FinetuneModel=Vicuna-13B, Sampling Ratio=100%2025.03 | 68 | — | — | — | — | |
| DPVR-LFModel Scale=13B, Split layer (s)=282026.06 | 67.6 | — | — | — | — | |
| DPVR-PCModel Scale=13B, Split layer (s)=282026.06 | 67.3 | — | — | — | — | |
| DPVR-PCModel Scale=13B, Split layer (s)=202026.06 | 67.2 | — | — | — | — | |
| DPVR-PCModel Scale=13B, Split layer (s)=242026.06 | 67.1 | — | — | — | — | |
| DPVR-PCModel Scale=13B, Split layer (s)=342026.06 | 67.1 | — | — | — | — | |
| LLaVoltaStages=Single, Scheme=no compression, #Tokens=18432, Compression Ratio (CR)=-, TFLOPS=8.26, Latency (ms)=68.5, Train Time=15.3h2024.06 | 66.7 | — | — | — | — | |
| VCDModel=Kimi-VL A3B-Thinking2025.10 | 66.52 | — | — | — | 49.69 | |
| OursModel=Ocean-R1 7B Instruct2025.10 | 66.51 | — | — | — | 49.66 | |
| R1-Onevision-7B2025.07 | 66.5 | — | — | — | — | |
| VanillaModel=Kimi-VL A3B-Thinking2025.10 | 66.26 | — | — | — | 49.54 | |
| LLaVoltaStages=Four, Scheme=deeper then wider, #Tokens=10863, Compression Ratio (CR)=170%, TFLOPS=8.26, Latency (ms)=68.5, Train Time=12.8h2024.06 | 66.1 | — | — | — | — | |
| LLaVoltaStages=Three, Scheme=compr. deeper, #Tokens=10597, Compression Ratio (CR)=174%, TFLOPS=8.26, Latency (ms)=68.5, Train Time=12.8h2024.06 | 65.9 | — | — | — | — | |
| LLaVoltaStages=Four, Scheme=wider then deeper, #Tokens=11088, Compression Ratio (CR)=166%, TFLOPS=8.26, Latency (ms)=68.5, Train Time=12.9h2024.06 | 65.6 | — | — | — | — | |
| LLaVoltaStages=Three, Scheme=last stage compression, #Tokens=7848, Compression Ratio (CR)=235%, TFLOPS=5.47, Latency (ms)=52.2, Train Time=12.4h2024.06 | 65.4 | — | — | — | — | |
| LLaVoltaStages=Three, Scheme=compr. wider, #Tokens=10407, Compression Ratio (CR)=177%, TFLOPS=8.26, Latency (ms)=68.5, Train Time=12.8h2024.06 | 65.3 | — | — | — | — | |
| RandomModel=Vicuna-13B, Sampling Ratio=15%2025.03 | 65 | — | — | — | — | |
| PreSelModel=Vicuna-13B, Sampling Ratio=15%2025.03 | 65 | — | — | — | — | |
| LLaVoltaStages=Two, Scheme=compression, #Tokens=10062, Compression Ratio (CR)=183%, TFLOPS=8.26, Latency (ms)=68.5, Train Time=12.8h2024.06 | 64.9 | — | — | — | — | |
| CGDModel=R1-Onevision 7B2025.10 | 63.23 | — | — | — | 46.58 | |
| AGLAModel=Ocean-R1 7B Instruct2025.10 | 63.01 | — | — | — | 45.95 | |
| SIMABase Model=LLaVA-1.5-13B2024.05 | 63 | — | — | — | — | |
| VILALLM=Llama 2-13B, Resolution=3362024.08 | 62.8 | — | — | — | — | |
| VILA-13BModel Scale=13B, Precision=FP162023.06 | 62.8 | — | — | — | — | |
| SIMABase Model=VILA-7B2024.05 | 62.5 | — | — | — | — | |
| GT-DPOBase Model=LLaVA-1.5-13B2024.05 | 62.2 | — | — | — | — | |
| VILA-13B-AWQModel Scale=13B, Quantization=INT4-g1282023.06 | 62.2 | — | — | — | — | |
| GT-DPOBase Model=VILA-7B2024.05 | 61.9 | — | — | — | — | |
| VILA-7BModel Scale=7B, Precision=FP162023.06 | 61.7 | — | — | — | — | |
| LLaVA-1.5LLM=Vicuna-1.5-13B, Resolution=3362024.08 | 61.6 | — | — | — | — | |
| LLaVA-1.5-13BBase Model=LLaVA-1.5-13B2024.05 | 61.6 | — | — | — | — | |
| mPLUG-Owl2LLM=LLaMA-7B2024.08 | 61.6 | — | — | — | — | |
| VILA-7B-AWQModel Scale=7B, Quantization=INT4-g1282023.06 | 61.3 | — | — | — | — | |
| OneLLM-7BLLM=LLAMA2-7B2023.12 | 61.2 | — | — | — | — | |
| VILALLM=Llama 2-7B, Resolution=3362024.08 | 61.1 | — | — | — | — | |
| VILA-7BBase Model=VILA-7B2024.05 | 61.1 | — | — | — | — | |
| SIMABase Model=LLaVA-1.5-7B2024.05 | 60.6 | — | — | — | — | |
| GT-DPOBase Model=LLaVA-1.5-7B2024.05 | 60.4 | — | — | — | — | |
| POVIDBase Model=LLaVA-1.5-7B2024.05 | 60.3 | — | — | — | — | |
| HA-DPOBase Model=LLaVA-1.5-7B2024.05 | 60.2 | — | — | — | — | |
| FASTLLM=Vicuna-7B2024.08 | 60.1 | — | — | — | — | |
| RLHFBase Model=LLaVA-1.5-7B2024.05 | 60 | — | — | — | — | |
| VCDModel=Ocean-R1 7B Instruct2025.10 | 59.98 | — | — | — | 42.87 | |
| VanillaModel=Ocean-R1 7B Instruct2025.10 | 59.76 | — | — | — | 42.61 | |
| Chain of SpotLLM=Vicuna-7B2024.08 | 59.7 | — | — | — | — | |
| LLaVA-1.5Model Size=7B, Backbone=Mistral2024.10 | 58.8 | 66.5 | 37.4 | 37.2 | — | |
| TGAModel Size=7B2024.10 | 58.7 | 66.3 | 37.4 | 37.7 | — | |
| LLaVA-1.5LLM=Vicuna-1.5-7B, Resolution=3362024.08 | 58.6 | — | — | — | — | |
| LLaVA-1.5-7BBase Model=LLaVA-1.5-7B2024.05 | 58.6 | — | — | — | — | |
| LLaVA-v1.5LLM=Vicuna-7B2023.12 | 58.6 | — | — | — | — | |
| LLaVA-v1.5LLM=Vicuna-7B2024.08 | 58.6 | — | — | — | — | |
| LLaVA-1.5Model Size=7B2024.10 | 58.6 | 66.1 | 37.3 | 37 | — | |
| InstructBLIPLLM=Vicuna-13B, Resolution=2242024.08 | 58.2 | — | — | — | — | |
| Qwen-VL-ChatLLM=Qwen-7B, Resolution=4482024.08 | 58.2 | — | — | — | — | |
| Qwen-VLLLM=Qwen-7B2023.12 | 58.2 | — | — | — | — | |
| Qwen-VL-ChatLLM=Qwen-7B2024.08 | 58.2 | — | — | — | — | |
| Qwen-VL-chatModel Size=7B2024.10 | 58.2 | 65.4 | 37.8 | 36.8 | — | |
| Qwen-VLLLM=Qwen-7B, Resolution=4482024.08 | 56.3 | — | — | — | — | |
| LLaVA-1.5-7BReq. Inst.=100%, Sel. Inst.=186K2025.03 | 55.6 | — | — | — | — | |
| COINCIDEReq. Inst.=100%, Sel. Inst.=28K2025.03 | 53.9 | — | — | — | — | |
| GPT-4V2025.07 | 53.8 | — | — | — | — | |
| PreSelReq. Inst.=15%, Sel. Inst.=28K2025.03 | 53.5 | — | — | — | — |