Multi-Task Reinforcement Learning (LTL Instruction Following) on Warehouse Infinite Horizon
880.6Average VisitsStructLTL
Evaluation Results
| Method | Links | |
|---|---|---|
| StructLTLformula=psi_11, zero-shot=true2026.02 | 880.6 | |
| StructLTLformula=psi_16, zero-shot=true2026.02 | 857.86 | |
| DeepLTLformula=psi_11, zero-shot=true2026.02 | 823.49 | |
| DeepLTLformula=psi_15, zero-shot=true2026.02 | 783.72 | |
| StructLTLformula=psi_13, zero-shot=true2026.02 | 682.62 | |
| StructLTLformula=psi_12, zero-shot=true2026.02 | 656.79 | |
| DeepLTLformula=psi_13, zero-shot=true2026.02 | 650.87 | |
| StructLTLformula=psi_15, zero-shot=true2026.02 | 577.07 | |
| DeepLTLformula=psi_16, zero-shot=true2026.02 | 575.43 | |
| DeepLTLformula=psi_12, zero-shot=true2026.02 | 433.9 | |
| StructLTLformula=psi_14, zero-shot=true2026.02 | 351.43 | |
| DeepLTLformula=psi_14, zero-shot=true2026.02 | 219.22 | |
| StructLTLformula=psi_8, zero-shot=true2026.02 | 15.92 | |
| DeepLTLformula=psi_8, zero-shot=true2026.02 | 13.84 | |
| StructLTLformula=psi_10, zero-shot=true2026.02 | 3.93 | |
| StructLTLformula=psi_9, zero-shot=true2026.02 | 3.19 | |
| StructLTLformula=psi_7, zero-shot=true2026.02 | 2.78 | |
| DeepLTLformula=psi_7, zero-shot=true2026.02 | 2.52 | |
| DeepLTLformula=psi_9, zero-shot=true2026.02 | 2.32 | |
| DeepLTLformula=psi_10, zero-shot=true2026.02 | 2.07 |