Natural Language Inference on SNLI (train)
99.7AccuracyLexicalized classifier
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Lexicalized classifier2015.09 | 99.7 | — | |
| Lexicalized classifier2016.03 | 99.7 | — | |
| Lexicalized Classifier2016.06 | 99.7 | — | |
| Classifier2015.12 | 99.7 | — | |
| Classifier with handcrafted features2016.07 | 99.7 | — | |
| Classifier with handcrafted features2016.07 | 99.7 | — | |
| + Unigram and bigram features2017.09 | 99.7 | — | |
| Bowman et al.features=handcrafted2018.02 | 99.7 | — | |
| Handcrafted features#Para.=-2016.09 | 99.7 | — | |
| 1024D pretrained GRU encodersParams.=15m2016.03 | 98.8 | — | |
| 1024D pretrained GRU encoders#Parameters=15.0M2016.06 | 98.8 | — | |
| 1024D GRU encoders|θ|=15m2017.09 | 98.8 | — | |
| Vendrov et al.2018.02 | 98.8 | — | |
| 1024D pretrained GRU encoders#Para.=15M2016.09 | 98.8 | — | |
| Human Performance (Estimated)2018.02 | 97.2 | — | |
| DR-BiLSTMmode=Ensemble2018.02 | 94.8 | — | |
| DR-BiLSTMmode=Ensemble, preprocessing=enabled2018.02 | 94.8 | — | |
| DR-BiLSTMmode=Single2018.02 | 94.1 | — | |
| DR-BiLSTMmode=Single, preprocessing=enabled2018.02 | 94.1 | — | |
| Chen et al.mode=Ensemble2018.02 | 93.5 | — | |
| HIM (600D ESIM + 300D Syntactic tree-LSTM)#Para.=7.7M2016.09 | 93.5 | — | |
| Wang et al.mode=Ensemble2018.02 | 93.2 | — | |
| 600D Gumbel TreeLSTM encoders|θ|=10m2018.01 | 93.1 | — | |
| 300D ReSAN|θ|=3.1m, T(s)/epoch=6222018.01 | 92.6 | — | |
| Chen et al.mode=Single2018.02 | 92.6 | — | |
| 600D ESIM#Para.=4.3M2016.09 | 92.6 | — | |
| Fine-tuningBackbone=RoBERTa-large, Training size=Full training set2021.08 | 92.6 | — | |
| Gong et al.mode=Ensemble2018.02 | 92.3 | — | |
| 300D mLSTM#Parameters=1.9M2016.06 | 92 | — | |
| mLSTMd=300, total number of parameters (|θ|_{W+M})=1.9M, parameters excluding word embeddings (|θ|_M)=1.9M2015.12 | 92 | — | |
| mLSTM word-by-word attentiond=300, number of parameters=1.9M2016.07 | 92 | — | |
| mLSTM word-by-word attentionword embedding size (d)=300, number of parameters (|theta|_M)=1.9M2016.07 | 92 | — | |
| Wang and Jiang2018.02 | 92 | — | |
| 300D mLSTM#Para.=1.9M2016.09 | 92 | — | |
| Bi-GRU|θ|=2.5m, T(s)/epoch=17282018.01 | 91.9 | — | |
| Hierarchical CNN|θ|=3.4m, T(s)/epoch=3432018.01 | 91.3 | — | |
| mLSTM with bi-LSTM sentence modelingd=150, total number of parameters (|θ|_{W+M})=1.4M, parameters excluding word embeddings (|θ|_M)=1.4M2015.12 | 91.3 | — | |
| Gong et al.mode=Single2018.02 | 91.2 | — | |
| DISAN|θ|=2.4m, T(s)/epoch=5872018.01 | 91.1 | — | |
| Directional self-attention network (DiSAN)|θ|=2.35m, T(s)/epoch=5872017.09 | 91.08 | — | |
| 600D Residual stacked encoders|θ|=29m2018.01 | 91 | — | |
| mLSTMd=150, total number of parameters (|θ|_{W+M})=544K, parameters excluding word embeddings (|θ|_M)=544K2015.12 | 91 | — | |
| Wang et al.mode=Single2018.02 | 90.9 | — | |
| Sha et al.2018.02 | 90.7 | — | |
| 300D re-read LSTM#Para.=2.0M2016.09 | 90.7 | — | |
| 600D Deep Gated Attn.|θ|=11.6m2018.01 | 90.5 | — | |
| Decomposable Attention Model (intra-sentence attention)#Parameters=582K2016.06 | 90.5 | — | |
| Decomposable attention modeld=200, number of parameters=582K2016.07 | 90.5 | — | |
| Decomposable Attention Modelword embedding size (d)=200, number of parameters (|theta|_M)=580K2016.07 | 90.5 | — | |
| Parikh et al.2018.02 | 90.5 | — | |
| Intra-sentence attention + 200D decomposable attention model#Para.=580K2016.09 | 90.5 | — | |
| Bi-LSTM|θ|=2.9m, T(s)/epoch=20802018.01 | 90.4 | — | |
| Bi-LSTM with s2t self-attention|θ|=2.88m, T(s)/epoch=20802017.09 | 90.39 | — | |
| LM-BFFBackbone=RoBERTa-large, Training size=Full training set, Demonstration usage=without demonstration2021.08 | 90.3 | — | |
| DISAN without directions|θ|=2.35m, T(s)/epoch=5922017.09 | 90.18 | — | |
| Multi-head|θ|=2.0m, T(s)/epoch=3452018.01 | 89.6 | — | |
| Multi-head with s2t self-attention|θ|=1.98m, T(s)/epoch=3452017.09 | 89.58 | — | |
| Decomposable Attention Model (vanilla)#Parameters=382K2016.06 | 89.5 | — | |
| LSTMN with deep attention fusiond=450, number of parameters=3.4M2016.07 | 89.5 | — | |
| 200D decomposable attention model#Para.=380K2016.09 | 89.5 | — | |
| DARTBackbone=RoBERTa-large, Training size=Full training set2021.08 | 89.5 | — | |
| Multi-window CNN|θ|=1.4m, T(s)/epoch=2842018.01 | 89.3 | — | |
| 300D SPINN-PI encoders|θ|=3.7m2018.01 | 89.2 | — | |
| 300D SPINN-PI (parsed input) encodersParams.=3.7m2016.03 | 89.2 | — | |
| 300D SPINN-PI encoders#Parameters=3.7M2016.06 | 89.2 | — | |
| SPINN-PI encodersd=300, number of parameters=3.7M2016.07 | 89.2 | — | |
| SPINN-PI encodersword embedding size (d)=300, number of parameters (|theta|_M)=3.5M2016.07 | 89.2 | — | |
| 300D SPINN-PI encoders|θ|=3.7m2017.09 | 89.2 | — | |
| Bowman et al.year=20162018.02 | 89.2 | — | |
| 300D SPINN-PI encoders#Para.=3.7M2016.09 | 89.2 | — | |
| mLSTM with word embeddingd=300, total number of parameters (|θ|_{W+M})=1.3M, parameters excluding word embeddings (|θ|_M)=1.3M2015.12 | 88.6 | — | |
| 300D btree-LSTM encoders#Para.=2.0M2016.09 | 88.6 | — | |
| 450D LSTMN with deep attention fusion#Parameters=3.4M2016.06 | 88.5 | — | |
| Full tree matching NTI-SLSTM-LSTM global attentiond=300, number of parameters=3.2M2016.07 | 88.5 | — | |
| LSTMN with deep attention fusionword embedding size (d)=450, number of parameters (|theta|_M)=3.4M2016.07 | 88.5 | — | |
| Full tree matching NTI-SLSTM-LSTM global attentionword embedding size (d)=300, number of parameters (|theta|_M)=3.2M2016.07 | 88.5 | — | |
| Liu et al.variant=2016a2018.02 | 88.5 | — | |
| Yu and Munkhdalaivariant=2017b2018.02 | 88.5 | — | |
| 450D LSTMN with deep attention fusion#Para.=3.4M2016.09 | 88.5 | — | |
| 300D NTI-SLSTM-LSTM#Para.=3.2M2016.09 | 88.5 | — | |
| NTI-SLSTM-LSTM node-by-node tree attentionword embedding size (d)=300, number of parameters (|theta|_M)=4.2M2016.07 | 88.1 | — | |
| NTI-SLSTM-LSTM node-by-node global attentionword embedding size (d)=300, number of parameters (|theta|_M)=4.2M2016.07 | 87.6 | — | |
| Tree matching NTI-SLSTM-LSTM global attentionword embedding size (d)=300, number of parameters (|theta|_M)=3.2M2016.07 | 87.6 | — | |
| Tree matching NTI-SLSTM-LSTM tree attentionword embedding size (d)=300, number of parameters (|theta|_M)=3.2M2016.07 | 87.3 | — | |
| 300D SPINN (unparsed input) encodersParams.=2.7m2016.03 | 87.2 | — | |
| MMA-NSEd=300, number of parameters=6.3M2016.07 | 87.1 | — | |
| MMA-NSE attentiond=300, number of parameters=6.5M2016.07 | 86.9 | — | |
| four stacked TC-LSTMsk=50, Number of parameters (|θ|M)=190K2016.05 | 86.7 | — | |
| Word-by-word attentionk=100, two-way=true, total_parameters=3.9M, parameters_no_word_embeddings=252k2015.09 | 86.6 | — | |
| Attentionk=100, two-way=true, total_parameters=3.9M, parameters_no_word_embeddings=242k2015.09 | 86.5 | — | |
| 600D Bi-LSTM encoders|θ|=2.0m2018.01 | 86.4 | — | |
| 600D Bi-LSTM encoders|θ|=2.0m2017.09 | 86.4 | — | |
| Word Embedding with s2t self-attention|θ|=0.54m, T(s)/epoch=2612017.09 | 86.22 | — | |
| 300D NSE encoders|θ|=3.0m2018.01 | 86.2 | — | |
| NSEd=300, number of parameters=3.4M2016.07 | 86.2 | — | |
| 300D NSE encoders|θ|=3.0m2017.09 | 86.2 | — | |
| Yu and Munkhdalaivariant=2017a2018.02 | 86.2 | — | |
| 300D NSE encoders#Para.=3.0M2016.09 | 86.2 | — | |
| NTI-SLSTM node-by-node tree attentionword embedding size (d)=300, number of parameters (|theta|_M)=3.5M2016.07 | 86 | — | |
| Word-by-word attention (our implementation)d=150, total number of parameters (|θ|_{W+M})=340K, parameters excluding word embeddings (|θ|_M)=340K2015.12 | 85.5 | — |