OCR-based Visual Question Answering on OCRVQA
86.51Mean AccuracyMetaForge
Evaluation Results
| Method | Links | |
|---|---|---|
| MetaForgeTool Setting=w/ IID Tools2026.06 | 86.51 | |
| MetaForgeTool Setting=w/ OOD Tools2026.06 | 86 | |
| MaLoRAModel=Qwen2.5-VL-7B, Training examples=801.6k2025.10 | 81.46 | |
| MaLoRAModel=Qwen3-VL-8B, Training examples=801.6k2025.10 | 81.39 | |
| LoRAModel=Qwen2.5-VL-7B, Training examples=801.6k2025.10 | 81.28 | |
| LoRAModel=Qwen3-VL-8B, Training examples=801.6k2025.10 | 81.02 | |
| GPT-5.4Tool Setting=w/ IID Tools2026.06 | 80.25 | |
| BaseModel=Qwen2.5-VL-7B, Training examples=801.6k2025.10 | 79.73 | |
| QLoRAModel=Qwen3-VL-8B, Training examples=801.6k2025.10 | 77.37 | |
| BaseModel=Qwen3-VL-8B, Training examples=801.6k2025.10 | 76.69 | |
| BaselineKeep ratio=100%, FLOPs Saved=—2026.05 | 73 | |
| QLoRAModel=Qwen2.5-VL-7B, Training examples=801.6k2025.10 | 72.98 | |
| SparseVLMKeep ratio=75%, FLOPs Saved=18%2026.05 | 72.8 | |
| FastVKeep ratio=75%, FLOPs Saved=0%2026.05 | 72.6 | |
| FastVKeep ratio=65%, FLOPs Saved=0%2026.05 | 72.6 | |
| SparseVLMKeep ratio=65%, FLOPs Saved=35%2026.05 | 72.6 | |
| FastVKeep ratio=50%, FLOPs Saved=0%2026.05 | 72.6 | |
| SparseVLMKeep ratio=50%, FLOPs Saved=47%2026.05 | 72.6 | |
| AsymVLMKeep ratio=75%, FLOPs Saved=28%2026.05 | 72.4 | |
| AsymVLMKeep ratio=65%, FLOPs Saved=39%2026.05 | 72 | |
| AsymVLMKeep ratio=50%, FLOPs Saved=54%2026.05 | 71.6 | |
| MaLoRAModel=LLaVA-1.5-7B, Training examples=801.6k2025.10 | 67.33 | |
| LoRAModel=LLaVA-1.5-7B, Training examples=801.6k2025.10 | 65.96 | |
| QLoRAModel=LLaVA-1.5-7B, Training examples=801.6k2025.10 | 65.62 | |
| Qwen3-VL-8B-InstructStrategy=Instruct2026.02 | 63.2 | |
| SAPStrategy=SAP2026.02 | 62.8 | |
| BaseModel=LLaVA-1.5-7B, Training examples=801.6k2025.10 | 60.99 | |
| Mixture Trainingmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 60.25 | |
| LLaVA-v1.5Backbone Parameters=7B2025.10 | 52.41 | |
| Iso-Cmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 46.51 | |
| OptMergemode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 46.35 | |
| WUDI Mergingmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 46.12 | |
| SWUDImode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 46.06 | |
| SWUDI-Amode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 45.9 | |
| TSV Mergingmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 45.38 | |
| KORErank=2352025.10 | 44.24 | |
| Qwen3-VL-8B-ThinkingStrategy=LongCoT2026.02 | 44.1 | |
| TIES Mergingmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 44.01 | |
| O-LoRA2025.10 | 43.91 | |
| Task Arithmeticmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 43.39 | |
| TIES w/ DAREmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 43.33 | |
| KORErank=2562025.10 | 43.03 | |
| TA w/ DAREmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 43.03 | |
| Weight Averagemode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 41.86 | |
| InternVL2.5-Instructmode=Full fine-tuning, backbone=InternVL2.5-1B2026.06 | 41.08 | |
| MoELoRA2025.10 | 40.17 | |
| Replay2025.10 | 37.73 | |
| CIA2025.10 | 34.31 | |
| SEFE2025.10 | 32.88 | |
| EWC2025.10 | 32.16 | |
| LwF2025.10 | 25.52 | |
| LoRA2025.10 | 23.8 | |
| (2)Adaptation strategy=Model parameter adaptation2025.10 | 13.8 | |
| TTAugAdaptation strategy=Test-time Augmentation2025.10 | 12.6 | |
| (1)Adaptation strategy=(1)2025.10 | 11.9 | |
| TTAugtest-time scaling=Method 52025.10 | 11.8 | |
| Full-FTLearning Strategy=Fine-Tuning2025.10 | 11.65 | |
| Method ④test-time scaling=Other method 42025.10 | 0.2 | |
| Baselinetest-time scaling=none2025.10 | 0 | |
| Method ①test-time scaling=Other method 12025.10 | 0 | |
| Method ②test-time scaling=Other method 22025.10 | 0 | |
| Method ③test-time scaling=Other method 32025.10 | 0 | |
| Baseline2025.10 | 0 |