Asynchronous Planning on NL (test)
78.2AccuracyGraph (40 steps) + NL (40 steps)
Evaluation Results
| Method | Links | |
|---|---|---|
| Graph (40 steps) + NL (40 steps)Model=Qwen 1.5B, Stage 1=Graph (40 steps), Stage 2=NL (40 steps), Training Algorithm=GRPO2026.02 | 78.2 | |
| GPT-4o (zero-shot)Mode=zero-shot2026.02 | 78.2 | |
| NL only (80 steps)Model=Qwen 1.5B, Total Training steps=80 steps, Training Algorithm=GRPO2026.02 | 69.8 | |
| 7B (NL 40 steps)Model=Qwen 7B, Training steps=40 steps2026.02 | 69.8 | |
| 3B (NL 40 steps)Model=Qwen 3B, Training steps=40 steps2026.02 | 47.1 | |
| GPT-4o-mini (zero-shot)Mode=zero-shot2026.02 | 44 | |
| NL (40 steps) + Graph (40 steps)Model=Qwen 1.5B, Stage 1=NL (40 steps), Stage 2=Graph (40 steps), Training Algorithm=GRPO2026.02 | 43.1 |