Puzzle Solving on Futoshiki (In Distribution)
79.8Avg@128Markov
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MarkovModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 79.8 | 96 | |
| MarkovModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 75 | 85 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 67.4 | 94 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 61.3 | 84 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 44.4 | 55 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 16.6 | 100 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 8.5 | 98 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 6.6 | 94 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0.3 | 20 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 0.1 | 7 | |
| MarkovModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0.1 | 11 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0 | 5 |