Puzzle Solving on Sudoku In Distribution
97.1Average Score @128Markov
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MarkovModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 97.1 | 98 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 93.5 | 97 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 91.1 | 96 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 86 | 94 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 83 | 90 | |
| MarkovModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 34.2 | 100 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 22.4 | 100 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 20 | 99 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 16.1 | 98 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 8.6 | 96 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0.3 | 23 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 0 | 0 |