Dialog act prediction on SwDA (test)
85Accuracyseq2seqBEST
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| seq2seqBESTencoder=HGRU, decoder=hard guided attention, Btrain=2, Binf=52020.02 | 85 | — | |
| Interlabeler agreement2016.03 | 84 | — | |
| Human annotator2017.11 | 84 | — | |
| Human Agreement2019.04 | 84 | — | |
| Human Agreement2018.10 | 84 | — | |
| Oursspeaker turn embeddings=true, topic-aware embeddings=false2021.09 | 83.2 | — | |
| SGNN2021.09 | 83.1 | — | |
| Our Method2019.04 | 82.9 | — | |
| SelfAtt-CRF2018.10 | 82.9 | — | |
| Li et al. (2018a)2020.02 | 82.9 | — | |
| Raheja and Tetreault (2019)2020.02 | 82.9 | — | |
| SelfAtt-CRF2021.09 | 82.9 | — | |
| Ours-Speakerspeaker turn embeddings=false, topic-aware embeddings=false2021.09 | 82.4 | — | |
| Ours+Topicspeaker turn embeddings=true, topic-aware embeddings=true2021.09 | 82.4 | — | |
| DAH-CRF + LDAutttopic labels=automatically acquired utterance-level via LDA2018.10 | 82.3 | — | |
| ALDMN2018.11 | 81.5 | — | |
| ALDMN2021.09 | 81.5 | — | |
| CRF-ASNmodel_type=deep neural network, structured prediction=true2017.11 | 81.3 | — | |
| Chen et al. (2018)2019.04 | 81.3 | — | |
| Chen et al. (2018)2020.02 | 81.3 | — | |
| DAH-CRF + MANUALconvtopic labels=manually annotated conversation-level2018.10 | 80.9 | — | |
| DAH-CRF-Manual2021.09 | 80.9 | — | |
| CRF-ASNreported_from=prior publication2018.10 | 80.8 | — | |
| CRF-ASN2021.09 | 80.8 | — | |
| DAH-CRF + LDAconvtopic labels=automatically acquired conversation-level via LDA2018.10 | 80.7 | — | |
| UCI2018.11 | 79.9 | — | |
| Li and Wu (2016)2019.04 | 79.4 | — | |
| Hierarchical Bi-LSTM-CRF2017.09 | 79.2 | — | |
| Bi-LSTM-CRFmodel_type=deep neural network, structured prediction=true2017.11 | 79.2 | — | |
| Kumar et al. (2018)2019.04 | 79.2 | — | |
| Bi-LSTM-CRF2018.10 | 79.2 | — | |
| Kumar et al. (2018)2020.02 | 79.2 | — | |
| Bi-LSTM-CRF2021.09 | 79.2 | — | |
| Average char-word-level & concatenated rep. predictionscontext=Utt-Att-BiRNN with context (WC)2018.05 | 77.42 | — | |
| Context-based RNNContext length=3 utterances2018.05 | 77.34 | 0.21 | |
| Context-based RNNContext length=4 utterances2018.05 | 77.28 | 0.22 | |
| DRLM-Conditional2017.09 | 77 | — | |
| DRLM-Conditionalmodel_type=deep neural network, structured prediction=true2017.11 | 77 | — | |
| DRLM-Conditional2018.11 | 77 | — | |
| Ji et al. (2016)2019.04 | 77 | — | |
| DRLM-Cond2018.10 | 77 | — | |
| DRLM-Cond2021.09 | 77 | — | |
| Average char-word-level predictionscontext=Utt-Att-BiRNN with context (WC)2018.05 | 76.84 | — | |
| Context-based RNNContext length=2 utterances2018.05 | 76.81 | 0.24 | |
| Context-based RNNContext length=1 utterance, SpeakerID=excluded2018.05 | 76.57 | 0.28 | |
| Context-based RNNContext length=1 utterance, SpeakerID=included2018.05 | 76.48 | 0.33 | |
| Character LM rep.context=Utt-Att-BiRNN with context (WC)2018.05 | 76.47 | — | |
| Concatenated rep.context=Utt-Att-BiRNN with context (WC)2018.05 | 76.15 | — | |
| LSTM-Softmax2017.09 | 75.8 | — | |
| LSTM-Softmaxmodel_type=deep neural network, structured prediction=false2017.11 | 75.8 | — | |
| BiLSTM-Softmax2018.11 | 75.8 | — | |
| Khanpour et al. (2016)2019.04 | 75.8 | — | |
| PDI2018.11 | 75.6 | — | |
| Word-embeddings mean rep.context=Utt-Att-BiRNN with context (WC)2018.05 | 75.43 | — | |
| DMN2018.11 | 75.2 | — | |
| Tran et al. (2017)2019.04 | 74.5 | — | |
| Proposed BaselineContext=Without context, Model type=RNN-based2018.05 | 73.96 | 0.26 | |
| Kalchbrenner and BlunsomModel type=Related previous work2018.05 | 73.9 | — | |
| RCNN2017.09 | 73.9 | — | |
| RCNNmodel_type=deep neural network, structured prediction=false2017.11 | 73.9 | — | |
| RCNN2018.11 | 73.9 | — | |
| Kalchbrenner and Blunsom (2013)2019.04 | 73.9 | — | |
| Lee and Dernoncourt (2016)2019.04 | 73.9 | — | |
| Kalchbrenner and Blunsomcontext=with context (WC)2018.05 | 73.9 | — | |
| Ortega and Vu (2017)2019.04 | 73.8 | — | |
| Ortega and Vucontext=with context (WC)2018.05 | 73.8 | — | |
| CNNselection=highest validation accuracy run2016.03 | 73.1 | — | |
| CNN2017.09 | 73.1 | — | |
| CNNmodel_type=deep neural network, structured prediction=false2017.11 | 73.1 | — | |
| Lee and Dernoncourtcontext=with context (WC)2018.05 | 73.1 | — | |
| Shen and Lee (2016)2019.04 | 72.6 | — | |
| Memory-based Learningfeatures=transcribed words and previous predicted dialog acts2016.03 | 72.3 | — | |
| CRFclassifier=CRF, features=pre-trained word embeddings2017.09 | 72.2 | — | |
| Average char-word-level & concatenated rep. predictionscontext=none (NC)2018.05 | 71.97 | — | |
| Average char-word-level predictionscontext=none (NC)2018.05 | 71.85 | — | |
| Character LM rep.context=none (NC)2018.05 | 71.84 | — | |
| Word-embeddings mean rep.context=none (NC)2018.05 | 71.73 | — | |
| CRFmodel_type=feature-based, structured prediction=true2017.11 | 71.7 | — | |
| LRclassifier=logistic regression, features=pre-trained word embeddings2017.09 | 71.4 | — | |
| JAS2018.10 | 71.2 | — | |
| HMMfeatures=transcribed words and previous predicted dialog acts2016.03 | 71 | — | |
| Stolcke et al.Model type=Related previous work2018.05 | 71 | — | |
| HMM2017.09 | 71 | — | |
| HMMmodel_type=feature-based, structured prediction=true2017.11 | 71 | — | |
| Stolcke et al.context=none (NC)2018.05 | 71 | — | |
| Concatenated rep.context=none (NC)2018.05 | 70.83 | — | |
| SVMmodel_type=feature-based, structured prediction=false2017.11 | 70.6 | — | |
| LSTMselection=highest validation accuracy run2016.03 | 69.6 | — | |
| TF-IDF GloVe2019.04 | 66.5 | — | |
| Majority class2016.03 | 33.7 | — | |
| Most common classModel type=Baseline2018.05 | 31.5 | — | |
| Most common class baselinecontext=none (NC)2018.05 | 31.5 | — |