Human-AI Planning Quality Assessment on 1,200 persona-goal profiles automatic simulation 1.0
4.27ScoreJumpStarter–Shallow
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| JumpStarter–Shallowmode=Shallow2024.10 | 4.27 | — | |
| JumpStarter–Recursivemode=Recursive2024.10 | 4.2 | 0.07 | |
| JumpStarter (No context elicitation)ablation=No context elicitation2024.10 | 4.19 | 0.08 | |
| JumpStarter (Random context selection)ablation=Random context selection2024.10 | 4.14 | 0.13 | |
| JumpStarter (No context reuse)ablation=No context reuse2024.10 | 4.09 | 0.18 | |
| JumpStarter (No context selection)ablation=No context selection2024.10 | 4.08 | 0.19 | |
| ADaPTtype=baseline2024.10 | 4.05 | 0.22 | |
| Ask-before-plantype=baseline2024.10 | 3.75 | 0.52 | |
| Memory-RAGtype=baseline2024.10 | 3.72 | 0.55 | |
| ChatGPT + elicited contexttype=baseline2024.10 | 3.29 | 0.98 | |
| ChatGPT + structured summarytype=baseline2024.10 | 3.29 | 0.98 | |
| ChatGPT vanillatype=baseline2024.10 | 3.22 | 1.05 |