Iterative Debugging on OR-Debug-Bench
95.3RR@58B model (RLVR training)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| 8B model (RLVR training)Model size=8B, Training Framework=RLVR with process rewards, Domain-specific Training=GRPO2026.01 | 95.3 | 62.4 | 2.25 | |
| Frontier APIsMethod Category=Frontier APIs2026.01 | 86.2 | 47.8 | 3.78 |