Graph prediction on CrossTask (test)
0.795Change Tire SuccessGround-truth labels → Graph
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Ground-truth labels → GraphInput source=Ground truth key step sequences2023.02 | 0.795 | 0.896 | 0.68 | 0.688 | 0.727 | 0.757 | |
| ASR → GPT → Labels → GraphInput source=Summary steps generated by GPT2023.02 | 0.75 | 0.729 | 0.617 | 0.625 | 0.659 | 0.676 | |
| Unsupervised Task Graph Generation PipelineInput source=Summary steps generated by GPT, Ranking/Filtering=Top-k filtering2023.02 | 0.727 | 0.771 | 0.617 | 0.625 | 0.682 | 0.684 | |
| ASR → Labels → GraphInput source=ASR sentences2023.02 | 0.545 | 0.729 | 0.711 | 0.562 | 0.545 | 0.618 | |
| ASR → VPs → Labels → GraphInput source=Verb phrases extracted from ASR2023.02 | 0.545 | 0.729 | 0.711 | 0.562 | 0.591 | 0.628 | |
| ProscriptModel Type=Fine-tuned language model, Training Data=Manually curated script data2023.02 | 0.523 | 0.896 | 0.57 | 0.625 | 0.614 | 0.646 |