Logical Reasoning on Zebra-CoT
25Jigsaw ScoreVisionR1
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VisionR1Paradigm=Reasoning, Fine-tuning Protocol=Official Checkpoints - No Task-specific Fine-tuning2025.12 | 25 | 45 | 65 | |
| Zero-shotParadigm=Direct Ans., Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset2025.12 | 23 | 44 | 65 | |
| ILVRParadigm=Interleaved, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 2: Latent Relaxation2025.12 | 22.5 | 47.8 | 73 | |
| CoT-FTParadigm=Text CoT, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset2025.12 | 21.5 | 45 | 68.5 | |
| ILVRParadigm=Interleaved, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 1: Latent Alignment2025.12 | 20.5 | 47.5 | 74.5 | |
| MirageParadigm=Single-step, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 2: Latent Relaxation2025.12 | 20 | 47.3 | 74.5 | |
| PixelReasonerParadigm=Tool-use, Fine-tuning Protocol=Official Checkpoints - No Task-specific Fine-tuning2025.12 | 18 | 45.5 | 73 | |
| Direct-FTParadigm=Direct Ans., Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset2025.12 | 17 | 45 | 73 | |
| MirageParadigm=Single-step, Fine-tuning Protocol=Fine-tuned on Zebra-CoT 10k subset, Training Stage=Stage 1: Latent Alignment2025.12 | 16 | 43.5 | 71 |