Semantic Parsing on OVERNIGHT v1.0 (test)
65.7Blocks Domain ScoreTWO-STAGE
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| TWO-STAGESupervision setting=Supervised2021.06 | 65.7 | 87.2 | 80.4 | 75.7 | 80.1 | 86.1 | 82.8 | 82.7 | 80.1 | |
| SSDSupervision setting=Supervised, Decoding granularity=Word-Level2021.06 | 64.9 | 86.2 | 81.7 | 72.7 | 82.3 | 81.7 | 81.5 | 82.7 | 79.2 | |
| SSDSupervision setting=Supervised, Decoding granularity=Grammar-Level2021.06 | 64.9 | 86.2 | 81.7 | 72.7 | 82.3 | 81.7 | 81.5 | 82.7 | 79 | |
| DUALSupervision setting=Supervised2021.06 | 63.7 | 87.5 | 79.8 | 73 | 81.4 | 81.5 | 81.6 | 83 | 78.9 | |
| GPT-3Representation=Canonical, Method=In-Context Learning, Decoding=Constrained2021.10 | 63.4 | 85.9 | 79.2 | 74.1 | 77.6 | 79.2 | 84 | 68.7 | 76.5 | |
| GPT-3Representation=Canonical, Method=In-Context Learning, Decoding=Constrained, evaluation_protocol=subsampled test set2021.10 | 62 | 80 | 82 | 71 | 79 | 84 | 89 | 72 | 77.4 | |
| T5-xlRepresentation=Canonical, Method=Prompt Tuning (PT), Decoding=Constrained2021.10 | 61.9 | 85.6 | 80.6 | 77.9 | 82.4 | 83 | 82.2 | 79.3 | 79.1 | |
| SEQ2ACTIONSupervision setting=Supervised2021.06 | 61.4 | 88.2 | 81.5 | 74.1 | 80.7 | 82.9 | 80.7 | 82.1 | 79 | |
| CROSSDOMAINSupervision setting=Supervised2021.06 | 60.2 | 86.2 | 79.8 | 71.4 | 78.9 | 84.7 | 81.6 | 82.9 | 78.2 | |
| T5-xlRepresentation=Meaning, Method=Prompt Tuning (PT), Decoding=Constrained2021.10 | 59.2 | 84.1 | 80.2 | 76.5 | 77.6 | 81.4 | 78.9 | 72.5 | 76.3 | |
| SSD-SAMPLESSupervision setting=Unsupervised (with nonparallel data), Decoding granularity=Grammar-Level2021.06 | 58.8 | 71.3 | 60.6 | 62.2 | 58.8 | 65.4 | 71.1 | 49.1 | 62.2 | |
| SSD-SAMPLESSupervision setting=Unsupervised (with nonparallel data), Decoding granularity=Word-Level2021.06 | 58.7 | 71.7 | 60.1 | 61.7 | 57.6 | 64.3 | 70.9 | 46 | 61.4 | |
| RECOMBINATIONSupervision setting=Supervised2021.06 | 58.1 | 85.2 | 78 | 71.4 | 76.4 | 79.6 | 76.2 | 81.4 | 75.8 | |
| SSDSupervision setting=Unsupervised, Decoding granularity=Grammar-Level2021.06 | 58.1 | 68.8 | 56.5 | 56.1 | 57.8 | 59.3 | 66.9 | 37.1 | 57.6 | |
| BARTRepresentation=Canonical, Method=Fine-Tuning (FT), Decoding=Constrained2021.10 | 55.4 | 86.4 | 78 | 67.2 | 75.8 | 80.1 | 80.1 | 66.6 | 73.7 | |
| SSDSupervision setting=Unsupervised, Decoding granularity=Word-Level2021.06 | 54.9 | 68.3 | 51.2 | 55 | 54.7 | 60.2 | 65.4 | 33.6 | 55.4 | |
| GPT-2Representation=Canonical, Method=Fine-Tuning (FT), Decoding=Constrained2021.10 | 54 | 83.6 | 76.6 | 66.6 | 71.5 | 76.4 | 76.8 | 62.3 | 71 | |
| TWO-STAGESupervision setting=Unsupervised (with nonparallel data)2021.06 | 53.4 | 64.7 | 58.3 | 59.3 | 60.3 | 68.1 | 73.2 | 48.4 | 60.7 | |
| GPT-3Representation=Meaning, Method=In-Context Learning, Decoding=Constrained, evaluation_protocol=subsampled test set2021.10 | 53 | 68 | 68 | 58 | 63 | 75 | 78 | 63 | 65.7 | |
| BARTRepresentation=Meaning, Method=Fine-Tuning (FT), Decoding=Constrained2021.10 | 49.9 | 83.4 | 75 | 61.9 | 73.9 | 79.6 | 77.4 | 62 | 70.4 | |
| GPT-2Representation=Meaning, Method=Fine-Tuning (FT), Decoding=Constrained2021.10 | 47.9 | 76 | 73.6 | 57.1 | 64.5 | 69.9 | 66 | 60.6 | 64.4 | |
| SYNTHPARA-SEQ2SEQSupervision setting=Unsupervised2021.06 | 37.3 | 28.4 | 33.9 | 38.1 | 39.1 | 41.7 | 62.7 | 23.3 | 38.1 | |
| WMDSAMPLESSupervision setting=Unsupervised (with nonparallel data)2021.06 | 29 | 31.9 | 36.1 | 47.9 | 34.2 | 41 | 53.8 | 35.8 | 38.7 | |
| Cross-domain Zero ShotSupervision setting=Unsupervised2021.06 | 28.3 | — | 53.6 | 52.4 | 55.3 | 60.2 | 61.7 | — | — | |
| GENOVERNIGHTSupervision setting=Unsupervised2021.06 | 27.7 | 15.6 | 17.3 | 45.9 | 46.7 | 26.3 | 61.3 | 9.7 | 31.3 | |
| SYNTH-SEQ2SEQSupervision setting=Unsupervised2021.06 | 23.6 | 16.1 | 16.1 | 30.2 | 36.6 | 26.9 | 43.1 | 9.2 | 25.2 |