Vision-Language Hallucination Evaluation on HallBench
64.2AccuracyOctopus-8B (Ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| Octopus-8B (Ours)Rollout count (rollout.n)=8, Generation Time (Gen.)=344.7, Total training time per step=958.12026.02 | 64.2 | |
| Qwen3-VL-8B-Instruct + DAPORollout count (rollout.n)=16, Learning Protocol=RLVR2026.02 | 63.7 | |
| MiMo-VL-7B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 63.5 | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=16, Generation Time (Gen.)=679.1, Total training time per step=1428.4, Learning Protocol=RLVR2026.02 | 62.8 | |
| Qwen3-VL-8B-ThinkingLearning Protocol=Open-source Reasoning VLM2026.02 | 62.7 | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=16, Generation Time (Gen.)=688.9, Total training time per step=1322.8, Learning Protocol=RLVR2026.02 | 62.5 | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=8, Generation Time (Gen.)=361.7, Total training time per step=753.1, Learning Protocol=RLVR2026.02 | 62.3 | |
| MiMo-VL-7B-SFTLearning Protocol=Open-source Reasoning VLM2026.02 | 62.1 | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=8, Generation Time (Gen.)=364.9, Total training time per step=845.1, Learning Protocol=RLVR2026.02 | 61.6 | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=16, Generation Time (Gen.)=816.6, Total training time per step=1543.3, Learning Protocol=RLVR2026.02 | 61.2 | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=8, Generation Time (Gen.)=410.9, Total training time per step=895.8, Learning Protocol=RLVR2026.02 | 60.8 | |
| Qwen3-VL-8B-InstructLearning Protocol=Base VLM2026.02 | 58.8 | |
| GPT-4oLearning Protocol=Closed-source VLM2026.02 | 56.2 | |
| Claude-3.7-SonnetLearning Protocol=Closed-source VLM2026.02 | 55.4 | |
| InternVL3.5-8B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 54.5 |