Chart-based Reasoning on CharXivRQ
67.9AccuracyGemini 2.5 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 2.5 Pro2026.04 | 67.9 | |
| Claude-3.7-SonnetLearning Protocol=Closed-source VLM2026.02 | 64.2 | |
| Octopus-8B (Ours)Rollout count (rollout.n)=8, Generation Time (Gen.)=344.7, Total training time per step=958.12026.02 | 55.7 | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=16, Generation Time (Gen.)=688.9, Total training time per step=1322.8, Learning Protocol=RLVR2026.02 | 55.3 | |
| OpenAI-o1Learning Protocol=Closed-source VLM2026.02 | 55.1 | |
| MiMo-VL-7B-SFTLearning Protocol=Open-source Reasoning VLM2026.02 | 54.8 | |
| MiMo-VL-7B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 53.2 | |
| Qwen3-VL-8B-ThinkingLearning Protocol=Open-source Reasoning VLM2026.02 | 53 | |
| OpenVLThinkerV2Parameters=8B2026.04 | 53 | |
| Qwen3-VL-8B-Instruct + DAPORollout count (rollout.n)=16, Learning Protocol=RLVR2026.02 | 52.8 | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=16, Generation Time (Gen.)=816.6, Total training time per step=1543.3, Learning Protocol=RLVR2026.02 | 52.7 | |
| Qwen3-VL GDPORL Strategy=GDPO2026.04 | 51.6 | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=16, Generation Time (Gen.)=679.1, Total training time per step=1428.4, Learning Protocol=RLVR2026.02 | 51.4 | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=8, Generation Time (Gen.)=410.9, Total training time per step=895.8, Learning Protocol=RLVR2026.02 | 51.2 | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=8, Generation Time (Gen.)=364.9, Total training time per step=845.1, Learning Protocol=RLVR2026.02 | 50.7 | |
| GPT-4oLearning Protocol=Closed-source VLM2026.02 | 50.5 | |
| Qwen3-VL GRPORL Strategy=GRPO2026.04 | 50.5 | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=8, Generation Time (Gen.)=361.7, Total training time per step=753.1, Learning Protocol=RLVR2026.02 | 47.9 | |
| GPT-4o2026.04 | 47.1 | |
| ARES-7BParameters=7B2026.04 | 47 | |
| Qwen3-VL-8B-InstructLearning Protocol=Base VLM2026.02 | 45.1 | |
| OVR-7BParameters=7B2026.04 | 44.5 | |
| Qwen3-VL-Instruct-8BParameters=8B2026.04 | 44.5 | |
| InternVL3.5-8B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 44.4 | |
| OneThinker-8BParameters=8B2026.04 | 44 | |
| Vision-G12026.04 | 41 | |
| VL-Rethinker-7BParameters=7B2026.04 | 39.8 | |
| MM-Eureka-7BParameters=7B2026.04 | 39.5 | |
| OpenVLThinker-7BParameters=7B2026.04 | 39.3 |