Multimodal Reasoning on MMStar (Acc., AvgLen, Ratio)
75.2AccuracyOctopus-8B (Ours)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Octopus-8B (Ours)Rollout count (rollout.n)=8, Generation Time (Gen.)=344.7, Total training time per step=958.12026.02 | 75.2 | — | — | |
| Qwen3-VL-8B-Instruct + DAPORollout count (rollout.n)=16, Learning Protocol=RLVR2026.02 | 74.7 | — | — | |
| Qwen3-VL-8B-ThinkingLearning Protocol=Open-source Reasoning VLM2026.02 | 74.1 | — | — | |
| MiMo-VL-7B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 73.7 | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=16, Generation Time (Gen.)=816.6, Total training time per step=1543.3, Learning Protocol=RLVR2026.02 | 73.3 | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=16, Generation Time (Gen.)=679.1, Total training time per step=1428.4, Learning Protocol=RLVR2026.02 | 73.1 | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=8, Generation Time (Gen.)=410.9, Total training time per step=895.8, Learning Protocol=RLVR2026.02 | 72.9 | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=16, Generation Time (Gen.)=688.9, Total training time per step=1322.8, Learning Protocol=RLVR2026.02 | 72.7 | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=8, Generation Time (Gen.)=361.7, Total training time per step=753.1, Learning Protocol=RLVR2026.02 | 72.5 | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=8, Generation Time (Gen.)=364.9, Total training time per step=845.1, Learning Protocol=RLVR2026.02 | 72.1 | — | — | |
| MiMo-VL-7B-SFTLearning Protocol=Open-source Reasoning VLM2026.02 | 70 | — | — | |
| Qwen3-VL-8B-InstructLearning Protocol=Base VLM2026.02 | 69.7 | — | — | |
| InternVL3.5-8B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 69.3 | — | — | |
| Baseline2026.02 | 68.25 | — | — | |
| Claude-3.7-SonnetLearning Protocol=Closed-source VLM2026.02 | 65.1 | — | — | |
| GPT-4oLearning Protocol=Closed-source VLM2026.02 | 64.7 | — | — | |
| XMCCBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 59 | 86.8 | 1.5 | |
| No CompressionBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 58.3 | 422.6 | 7.2 | |
| StepEntropyBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 58.1 | 225.5 | 3.9 | |
| IDPrunerRetain Tokens=25%2026.02 | 57.97 | — | — | |
| Prune-on-LogicBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 57.7 | 229 | 4 | |
| VisionSelectorRetain Tokens=25%2026.02 | 57.42 | — | — | |
| DARTRetain Tokens=25%2026.02 | 56.8 | — | — | |
| XMCCBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 56.4 | 87.5 | 1.6 | |
| Prune-on-LogicBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 56.2 | 248 | 4.4 | |
| No CompressionBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 55.1 | 452.7 | 8.2 | |
| StepEntropyBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 54.1 | 211.1 | 3.9 | |
| VisionSelectorRetain Tokens=10%2026.02 | 50.74 | — | — | |
| IDPrunerRetain Tokens=10%2026.02 | 50.02 | — | — |