Puzzle Solving on Sudoku Out Of Distribution
77.8Average @128Markov
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MarkovModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 77.8 | 82 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 71.2 | 82 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 69.2 | 82 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 57.1 | 69 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 56.4 | 68 | |
| MarkovModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 8.7 | 86 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 3.1 | 64 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 2.9 | 75 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 2.4 | 71 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 1.9 | 62 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0 | 0 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 0 | 0 |