Generalized Planning on Gripper
100ScaleState-value policy (V)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| State-value policy (V)Architecture=R-GNN2026.03 | 100 | 77.5 | |
| Explicitly regularized Q-value policy (Ω^Exp.)Architecture=R-GNN2026.03 | 100 | 88.8 | |
| Heuristically regularized Q-value policy (Ω^Heu.)Architecture=R-GNN2026.03 | 100 | 88.9 | |
| State-value policy (V)Architecture=OE2026.03 | 100 | 83 | |
| Explicitly regularized Q-value policy (Ω^Exp.)Architecture=OE2026.03 | 100 | 88.9 | |
| Heuristically regularized Q-value policy (Ω^Heu.)Architecture=OE2026.03 | 100 | 88.4 | |
| State-value policy (V)Architecture=OAE2026.03 | 100 | 88.6 | |
| Vanilla Q-value policy (Q)Architecture=OAE2026.03 | 100 | 73.8 | |
| Explicitly regularized Q-value policy (Ω^Exp.)Architecture=OAE2026.03 | 100 | 88.5 | |
| Heuristically regularized Q-value policy (Ω^Heu.)Architecture=OAE2026.03 | 100 | 88.6 | |
| Vanilla Q-value policy (Q)Architecture=OE2026.03 | 63 | 33.4 | |
| Vanilla Q-value policy (Q)Architecture=R-GNN2026.03 | 36 | 15.2 |