Multimodal Understanding on CVBench
64.25AccuracyLaViDa-O + RL (Global + Factorized)
Evaluation Results
| Method | Links | |
|---|---|---|
| LaViDa-O + RL (Global + Factorized)Arch=D-Diff, Prompting Strategy=CoT2026.06 | 64.25 | |
| LaViDa-OArch=D-Diff, Prompting Strategy=CoT2026.06 | 59.7 | |
| LaViDa-O + RL (LocFac-RL)Arch=D-Diff, Prompting Strategy=CoT, per-step train time reduction=↓ 12.3% s/step2026.06 | 51.9 | |
| LaViDa-O + SFTArch=D-Diff, Prompting Strategy=CoT2026.06 | 48.67 | |
| MMaDA-Parallel + RL (LocFac-RL)Arch=D-Diff, Prompting Strategy=CoT, per-step train time reduction=↓ 16.9% s/step2026.06 | 37.98 | |
| MMaDA-Parallel + SFTArch=D-Diff, Prompting Strategy=CoT2026.06 | 36.81 | |
| MMaDA-Parallel + RL (Global + Factorized)Arch=D-Diff, Prompting Strategy=CoT2026.06 | 36.54 | |
| ANOLE + RLArch=AR, Prompting Strategy=CoT2026.06 | 18.4 | |
| ANOLE + SFTArch=AR, Prompting Strategy=CoT2026.06 | 17.6 | |
| ANOLEArch=AR, Prompting Strategy=CoT2026.06 | 14.4 | |
| MMaDA-ParallelArch=D-Diff, Prompting Strategy=CoT2026.06 | 3.87 |