Puzzle Solving on Futoshiki (Out Of Distribution)
42.6Avg@128Markov
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MarkovModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 42.6 | 53 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 28.3 | 67 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 25.2 | 75 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=RL post-training2026.03 | 24.9 | 56 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 16.9 | 21 | |
| State-action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 1.1 | 60 | |
| Action-sequenceModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0.3 | 28 | |
| MarkovModel=Qwen2.5-3B-It, Training Stage=SFT Warm-up2026.03 | 0.3 | 26 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0 | 0 | |
| Action-sequenceModel=Qwen3-4B, Training Stage=RL post-training2026.03 | 0 | 0 | |
| MarkovModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0 | 0 | |
| State-action-sequenceModel=Qwen3-4B, Training Stage=SFT Warm-up2026.03 | 0 | 1 |