Natural Language Inference on SNLI (dev)
93.6AccuracyALUM_ROBERTA-LARGE-SMART
Evaluation Results
| Method | Links | |
|---|---|---|
| ALUM_ROBERTA-LARGE-SMARTBackbone=RoBERTa-Large, Adversarial Pre-training=ALUM, Adversarial Fine-tuning=SMART2020.04 | 93.6 | |
| ALUM_ROBERTA-LARGEBackbone=RoBERTa-Large, Adversarial Pre-training=ALUM2020.04 | 93.1 | |
| EFLTraining Setting=Full training dataset, Backbone=RoBERTa-large2021.04 | 93.1 | |
| MT-DNN-SMART_LARGE_V0backbone=BERT-Large, architecture=MT-DNN, framework=SMART, version=v02019.11 | 92.6 | |
| MT-DNN_LARGEbackbone=BERT-Large, architecture=MT-DNN2019.11 | 92.2 | |
| MT-DNNModel Architecture=LARGE2019.01 | 92.2 | |
| StructBERTLarge2019.08 | 92.2 | |
| MT-DNNevaluation protocol=fine-tuned2019.09 | 92.2 | |
| SemBERT_WWMscale=Large, pre-training=Whole Word Masking, evaluation protocol=fine-tuned2019.09 | 92.2 | |
| MT-DNN_LARGEBackbone=MT-DNN-Large2020.04 | 92.2 | |
| BERT_WWMscale=Large, pre-training=Whole Word Masking, evaluation protocol=fine-tuned2019.09 | 92.1 | |
| SemBERT_LARGEscale=Large, evaluation protocol=fine-tuned2019.09 | 92 | |
| MT-DNN-SMART_BASE_V0backbone=BERT-Base, architecture=MT-DNN, framework=SMART, version=v02019.11 | 91.7 | |
| MT-DNN-SMART_BASEbackbone=BERT-Base, architecture=MT-DNN, framework=SMART2019.11 | 91.7 | |
| BERT_LARGEbackbone=BERT-Large2019.11 | 91.7 | |
| BERTModel Architecture=LARGE2019.01 | 91.7 | |
| BERT_LARGEBackbone=BERT-Large2020.04 | 91.7 | |
| MT-DNNModel Architecture=BASE2019.01 | 91.5 | |
| MT-DNN_BASEbackbone=BERT-Base, architecture=MT-DNN2019.11 | 91.4 | |
| SMART_BERT-BASEbackbone=BERT-Base, framework=SMART2019.11 | 91.4 | |
| MT-DNN2019.08 | 91.4 | |
| BERT_LARGEscale=Large, evaluation protocol=fine-tuned2019.09 | 91.3 | |
| SemBERT_BASEscale=Base, evaluation protocol=fine-tuned2019.09 | 91.2 | |
| BERT teacherRole=Teacher, Architecture=BERT, Parameters=110 million2023.05 | 91.03 | |
| BERT_BASEbackbone=BERT-Base2019.11 | 91 | |
| BERTModel Architecture=BASE2019.01 | 91 | |
| BERT_BASEscale=Base, evaluation protocol=fine-tuned2019.09 | 90.8 | |
| BERT2019.08 | 90.1 | |
| ESIM GloVeAuxiliary Information=GloVe vectors2017.06 | 87.2 | |
| mLSTMd=300, total number of parameters (|θ|_{W+M})=1.9M, parameters excluding word embeddings (|θ|_M)=1.9M2015.12 | 86.9 | |
| mLSTM with bi-LSTM sentence modelingd=150, total number of parameters (|θ|_{W+M})=1.4M, parameters excluding word embeddings (|θ|_M)=1.4M2015.12 | 86.6 | |
| mLSTMd=150, total number of parameters (|θ|_{W+M})=544K, parameters excluding word embeddings (|θ|_M)=544K2015.12 | 86.2 | |
| mLSTM with word embeddingd=300, total number of parameters (|θ|_{W+M})=1.3M, parameters excluding word embeddings (|θ|_M)=1.3M2015.12 | 85.4 | |
| ESIM dictionaryAuxiliary Information=dictionary definitions2017.06 | 84.88 | |
| ESIM spellingAuxiliary Information=spelling2017.06 | 83.78 | |
| Word-by-word attentionk=100, two-way=false, total_parameters=3.9M, parameters_no_word_embeddings=252k2015.09 | 83.7 | |
| Word-by-word attentiond=100, total number of parameters (|θ|_{W+M})=3.9M, parameters excluding word embeddings (|θ|_M)=252K2015.12 | 83.7 | |
| Word-by-word attentionk=100, two-way=true, total_parameters=3.9M, parameters_no_word_embeddings=252k2015.09 | 83.6 | |
| ESIM baselineAuxiliary Information=none2017.06 | 83.39 | |
| Word-by-word attention (our implementation)d=150, total number of parameters (|θ|_{W+M})=340K, parameters excluding word embeddings (|θ|_M)=340K2015.12 | 83.3 | |
| Attentionk=100, two-way=false, total_parameters=3.9M, parameters_no_word_embeddings=242k2015.09 | 83.2 | |
| Conditional encodingk=159, shared=true, total_parameters=3.9M, parameters_no_word_embeddings=252k2015.09 | 83 | |
| Attentionk=100, two-way=true, total_parameters=3.9M, parameters_no_word_embeddings=242k2015.09 | 83 | |
| LSTM sharedd=159, total number of parameters (|θ|_{W+M})=3.9M, parameters excluding word embeddings (|θ|_M)=252K2015.12 | 83 | |
| ERNIE 2.0 (L)Backbone Scale=Large2019.11 | 82.6 | |
| NEZHA-wwm (L)Whole Word Masking=true, Backbone Scale=Large2019.11 | 82.21 | |
| Conditional encodingk=116, shared=false, total_parameters=3.9M, parameters_no_word_embeddings=252k2015.09 | 82.1 | |
| Conditional encodingk=100, shared=true, total_parameters=3.8M, parameters_no_word_embeddings=111k2015.09 | 81.9 | |
| NEZHA (L)Backbone Scale=Large2019.11 | 81.53 | |
| NEZHA (B)Backbone Scale=Base2019.11 | 81.37 | |
| NEZHA-WWM (B)Whole Word Masking=true, Backbone Scale=Base2019.11 | 81.25 | |
| ERNIE 2.0 (B)Backbone Scale=Base2019.11 | 81.2 | |
| EFLTraining Setting=Few-shot with K=8, Backbone=RoBERTa-large2021.04 | 81 | |
| ZEN (P)Initialization=Pre-trained, Backbone Scale=Base2019.11 | 80.48 | |
| DTARole=Student, Architecture=TinyBERT, Strategy=Domain-Targeted Augmentation2023.05 | 80.16 | |
| KD (standard distillation)Role=Student, Architecture=TinyBERT, Distillation Strategy=Standard KD2023.05 | 80.11 | |
| DTA with DMURole=Student, Architecture=TinyBERT, Strategy=DTA + DMU combined2023.05 | 80.11 | |
| KD + SmoothingRole=Student, Architecture=TinyBERT, Distillation Strategy=KD with label smoothing2023.05 | 80.09 | |
| DMURole=Student, Architecture=TinyBERT, Strategy=Distilled Minority Upsampling2023.05 | 80 | |
| ERNIE 1.0Backbone Scale=Base2019.11 | 79.9 | |
| BERT-WWMWhole Word Masking=true, Backbone Scale=Base2019.11 | 78.4 | |
| TinyBERT baselineRole=Student, Architecture=TinyBERT, Parameters=4.4 million, Distillation=None2023.05 | 77.99 | |
| Baseline w/ labelled aug. dataRole=Student, Architecture=TinyBERT, Data Augmentation=Labelled Hypothesis Generation2023.05 | 77.72 | |
| Fine-tuningTraining Setting=Full training dataset, Backbone=RoBERTa-large2021.04 | 77.6 | |
| BERT (P)Initialization=Pre-trained, Backbone Scale=Base2019.11 | 77.4 | |
| ZEN (R)Initialization=Random, Backbone Scale=Base2019.11 | 77.11 | |
| JTTRole=Student, Architecture=TinyBERT, Strategy=Just Train Twice2023.05 | 76.96 | |
| BERT (R)Initialization=Random, Backbone Scale=Base2019.11 | 75.67 | |
| LM-BFFTraining Setting=Few-shot with K=8, Backbone=RoBERTa-large2021.04 | 52 | |
| Fine-tuningTraining Setting=Few-shot with K=8, Backbone=RoBERTa-large2021.04 | 38.4 | |
| MajorityTraining Setting=Full training dataset, Backbone=RoBERTa-large2021.04 | 33.8 |