Visual Reasoning on HalluBench
71.85AccuracySaEI
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SaEIFinetuned=true, Group Size=122025.12 | 71.85 | — | |
| NoisyRolloutFinetuned=true, Group Size=122025.12 | 71.33 | — | |
| Shuffle-R1-Qwen-7BTraining Strategy=Zero RL2025.08 | 71 | — | |
| KL-CovFinetuned=true, Group Size=122025.12 | 70.67 | — | |
| Vanilla GRPOFinetuned=true, Group Size=122025.12 | 70.48 | — | |
| ThinkLite-VL-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 70.2 | — | |
| NoisyRollout-7B-K12Training Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 70.1 | — | |
| VL-Rethinker-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 69.9 | — | |
| MMR1-Math-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 69.6 | — | |
| Shuffle-R1-Qwen-3BTraining Strategy=Zero RL2025.08 | 69.2 | — | |
| VLAA-Thinker-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 67.5 | — | |
| MM-Eureka-Qwen-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 66.7 | — | |
| R1-OneVision-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 66.4 | — | |
| Qwen2.5-VL-7BTraining Strategy=Open-Source SFT, Evaluation Toolkit=vLLM with custom scripts2025.08 | 65.2 | — | |
| Qwen2.5-VL-7B-InstructFinetuned=false2025.12 | 64 | — | |
| R1-VL-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 60.9 | — | |
| Qwen2.5-VL-3BTraining Strategy=Open-Source SFT, Evaluation Toolkit=vLLM with custom scripts2025.08 | 59.8 | — | |
| OpenVLThinker-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 59.1 | — | |
| Vision-R1-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 57.8 | — | |
| Claude-3.7-SonnetTraining Strategy=Close-source2025.08 | 55.4 | — | |
| GPT-4oTraining Strategy=Close-source2025.08 | 55 | — | |
| InternVL-2.5-8BTraining Strategy=Open-Source SFT2025.08 | 50.1 | — | |
| InternVL-3-8BTraining Strategy=Open-Source SFT2025.08 | 49.9 | — | |
| Gemini-2.0 proTraining Strategy=Close-source2025.08 | 49.8 | — | |
| KL-CovFinetuning status=Finetuned, Group size=n = 122025.12 | — | 68.31 | |
| NoisyRolloutFinetuning status=Finetuned, Group size=n1 = n2 = 6 (total n = 12)2025.12 | — | 70.11 | |
| Qwen2.5-VL-7B-InstructFinetuning status=Not finetuned2025.12 | — | 64 | |
| SaEIFinetuning status=Finetuned, Group size=n1 = n2 = 6 (total n = 12)2025.12 | — | 70.38 | |
| Vanilla GRPOFinetuning status=Finetuned, Group size=n = 122025.12 | — | 70.18 |