Sudoku (Accuracy)
91.82AccuracyGlauber-UL2 (N=3)
Evaluation Results
| Method | Links | |
|---|---|---|
| Glauber-UL2 (N=3)Params=7M2026.05 | 91.82 | |
| MDLM (Top-K Margin)Params=6M2026.05 | 89.49 | |
| AR (w/ ordering)Params=42M2026.05 | 87.18 | |
| Glauber-UL2 (N=1)Params=7M2026.05 | 42.26 | |
| Qwen3-VL + V-ABSReasoning Strategy=V-ABS2026.05 | 37.8 | |
| Intern-VL3 + V-ABSReasoning Strategy=V-ABS2026.05 | 33.1 | |
| Qwen2.5-VL + V-ABSReasoning Strategy=V-ABS2026.05 | 32.5 | |
| Instruct Model + SFTShots=5-shot, Training Dataset=s1.1K, Tuning Method=LoRA Tuning2025.09 | 23.3 | |
| Instruct Model + GIFTShots=5-shot, Training Dataset=s1.1K, Tuning Method=LoRA Tuning2025.09 | 20.1 | |
| MDLM (Top-K Prob.)Params=6M2026.05 | 18.51 | |
| Qwen3-VL2026.05 | 14.4 | |
| Qwen2.5-VL2026.05 | 14 | |
| Intern-VL32026.05 | 11.3 | |
| AR (w/o ordering)Params=42M2026.05 | 9.73 | |
| Instruct Model + GIFTShots=5-shot, Training Dataset=s1K, Tuning Method=LoRA Tuning2025.09 | 7.9 | |
| MDLM (vanilla)Params=6M2026.05 | 6.88 | |
| Instruct Model + SFTShots=5-shot, Training Dataset=s1K, Tuning Method=LoRA Tuning2025.09 | 6 | |
| SFTBackbone Model=Dream-v0-7B-Base, Training Dataset=Tulu3-10k, Fine-tuning Protocol=Full-parameter Tuning, Evaluation Protocol=0-shot2025.09 | 3.8 | |
| GIFTBackbone Model=Dream-v0-7B-Base, Training Dataset=Tulu3-10k, Fine-tuning Protocol=Full-parameter Tuning, Evaluation Protocol=0-shot2025.09 | 2.1 |