Multimodal Reasoning on BLINK Jigsaw
55.33AccuracyMMaDA-Parallel + RL (LocFac-RL)
Evaluation Results
| Method | Links | |
|---|---|---|
| MMaDA-Parallel + RL (LocFac-RL)Arch=D-Diff, Prompting Strategy=CoT, per-step train time reduction=↓ 16.9% s/step2026.06 | 55.33 | |
| MMaDA-Parallel + RL (Global + Factorized)Arch=D-Diff, Prompting Strategy=CoT2026.06 | 52 | |
| MMaDA-Parallel + SFTArch=D-Diff, Prompting Strategy=CoT2026.06 | 47.33 | |
| LaViDa-O + RL (LocFac-RL)Arch=D-Diff, Prompting Strategy=CoT, per-step train time reduction=↓ 12.3% s/step2026.06 | 45.43 | |
| LaViDa-OArch=D-Diff, Prompting Strategy=CoT2026.06 | 44.77 | |
| LaViDa-O + SFTArch=D-Diff, Prompting Strategy=CoT2026.06 | 43.61 | |
| LaViDa-O + RL (Global + Factorized)Arch=D-Diff, Prompting Strategy=CoT2026.06 | 43.48 | |
| ANOLE + RLArch=AR, Prompting Strategy=CoT2026.06 | 38.7 | |
| ANOLE + SFTArch=AR, Prompting Strategy=CoT2026.06 | 36.7 | |
| ANOLEArch=AR, Prompting Strategy=CoT2026.06 | 28 | |
| MMaDA-ParallelArch=D-Diff, Prompting Strategy=CoT2026.06 | 0.68 |