Puzzle Solving on Sokoban (In Distribution)
89.7Average Score @128Markov
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MarkovModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 89.7 | 93 | |
| MarkovModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 76.1 | 81 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 57.4 | 67 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 43.6 | 50 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 2.3 | 4 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 1.4 | 61 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 1 | 1 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0.6 | 41 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0.5 | 37 | |
| MarkovModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0.4 | 28 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0.2 | 22 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0.2 | 14 |