Puzzle Solving on Sokoban (Out Of Distribution)
66.9Avg@128Markov
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MarkovModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 66.9 | 72 | |
| MarkovModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 31.6 | 37 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 30.2 | 34 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 20.4 | 23 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0.2 | 15 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0 | 1 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 0 | 0 | |
| MarkovModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0 | 3 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0 | 3 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 0 | 0 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0 | 1 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0 | 1 |