Visual Reasoning on VisuLogic
29.3Avg ScoreILVR
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ILVRParadigm=Interleaved, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 2: Latent Relaxation2025.12 | 29.3 | 27 | 30 | 31 | |
| Bagel-ZebraParadigm=Unified, Fine-tuning Protocol=Official Checkpoints - No Task-specific Fine-tuning2025.12 | 28.9 | 28 | 39 | 21 | |
| Zero-shotParadigm=Direct Ans., Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset2025.12 | 26.6 | 29 | 24 | 27 | |
| MirageParadigm=Single-step, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 2: Latent Relaxation2025.12 | 26.6 | 24 | 26 | 30 | |
| CoT-FTParadigm=Text CoT, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset2025.12 | 25.9 | 27 | 23 | 28 | |
| DeepLatent-RL-7BModel Category=Latent Visual reasoning Models, Training Stage=RL, Training Data=full DeepLatent-180K dataset2026.05 | 25.3 | — | — | — | |
| InternVL3-8BModel Category=Open-source Models2026.05 | 24.9 | — | — | — | |
| DeepLatent-SFT-7BModel Category=Latent Visual reasoning Models, Training Stage=SFT2026.05 | 24.8 | — | — | — | |
| ILVRParadigm=Interleaved, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 1: Latent Alignment2025.12 | 24.5 | 26 | 23 | 24 | |
| DeepLatent-RL-7B*Model Category=Latent Visual reasoning Models, Training Stage=RL, Training Data=visual search data of DeepLatent-180K in SFT Stage 22026.05 | 24.5 | — | — | — | |
| Direct-FTParadigm=Direct Ans., Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset2025.12 | 23.8 | 25 | 23 | 23 | |
| PixelReasonerParadigm=Tool-use, Fine-tuning Protocol=Official Checkpoints - No Task-specific Fine-tuning2025.12 | 23.4 | 18 | 16 | 29 | |
| MirageParadigm=Single-step, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 1: Latent Alignment2025.12 | 23.4 | 25 | 24 | 21 | |
| Thyme-7BModel Category=Tool-based Models2026.05 | 23.4 | — | — | — | |
| Qwen2.5-VL-7BModel Category=Open-source Models2026.05 | 20 | — | — | — | |
| VisionR1Paradigm=Reasoning, Fine-tuning Protocol=Official Checkpoints - No Task-specific Fine-tuning2025.12 | 15.2 | 18 | 13 | 14 |