Sudoku Solving on 9x9 Sudoku (test)
52Cell AccuracyFine-tuned (solver order)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Fine-tuned (solver order)Training Protocol=Supervised Fine-tuning2025.12 | 52 | — | — | |
| GRPO with Bootstrapped Mixed RewardsCell : Order Weight=0.75 : 0.25, Base Model=Fine-tuned (random order)2025.12 | 49.6 | — | — | |
| GRPO with Bootstrapped Mixed RewardsCell : Order Weight=0.5 : 0.5, Base Model=Fine-tuned (random order)2025.12 | 46.5 | — | — | |
| GRPO with Bootstrapped Mixed RewardsCell : Order Weight=1 : 0, Base Model=Fine-tuned (random order)2025.12 | 43.8 | — | — | |
| GRPO with Bootstrapped Mixed RewardsCell : Order Weight=0.25 : 0.75, Base Model=Fine-tuned (random order)2025.12 | 37.8 | — | — | |
| Fine-tuned (random order)Training Protocol=Supervised Fine-tuning2025.12 | 28.2 | — | — | |
| GRPO with Bootstrapped Mixed RewardsCell : Order Weight=0 : 1, Base Model=Fine-tuned (random order)2025.12 | 26.2 | — | — | |
| GPT-OSS-20BEvaluation protocol=Without task-specific program synthesis or symbolic search2026.03 | — | 21.67 | — | |
| HRMTraining Data=9x9 Sudoku, Inference Steps=162026.03 | — | 63.53 | 86.11 | |
| SE-RRMTraining Data=9x9 Sudoku, Inference Steps=162026.03 | — | 93.73 | 97.58 | |
| TRMTraining Data=9x9 Sudoku, Inference Steps=162026.03 | — | 71.94 | 89.8 |