Natural Language Inference on MNLI (matched)
91.7AccuracyDeBERTaxLarge
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeBERTaxLarge# Param.=900M2021.09 | 91.7 | — | |
| DeBERTaxxLarge# Param.=1.5B2021.09 | 91.7 | — | |
| METALMtuning_mode=Single-task finetuning, parameters_updated=non-causal encoder and connector, frozen_components=language model2022.06 | 91.1 | — | |
| ELECTRAtuning_mode=Single-task finetuning, model_size=Large2022.06 | 90.9 | — | |
| ALBERTxxLarge# Param.=223M2021.09 | 90.8 | — | |
| FLAN-T5Model Variant=xlarge, Model Parameters=3B2023.07 | 90.5 | — | |
| ALIGNModel Variant=large, Model Parameters=355M2023.07 | 90.3 | — | |
| RoBERTatuning_mode=Single-task finetuning, model_size=Large2022.06 | 90.2 | — | |
| RoBERTa# Param.=355M2021.09 | 90.2 | — | |
| BART# Param.=406M2021.09 | 89.9 | — | |
| RoBERTa-large2023.05 | 89.6 | — | |
| FLAN-T5Model Variant=large, Model Parameters=780M2023.07 | 88.8 | — | |
| ALIGNModel Variant=base, Model Parameters=125M2023.07 | 87.8 | — | |
| GPTtuning_mode=Single-task finetuning, parameters_updated=all2022.06 | 87.7 | — | |
| ComKD#aug=393k2023.05 | 87.2 | — | |
| BERTtuning_mode=Single-task finetuning, model_size=Large2022.06 | 86.6 | — | |
| Prompt-based FT (hard) + PCPBackbone=RoBERTa-LARGE, Supervision Setting=Full, Evaluation Protocol=Prompt-based Fine-tuning (hard), Continued Pre-training=PCP2023.05 | 86.5 | — | |
| DTAensBackbone=DeBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 85.77 | 1.21 | |
| KDBackbone=DeBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 84.83 | 0.27 | |
| CAMEmode=fine-tuning, batch size=8k2023.07 | 84.8 | — | |
| CAMEmode=fine-tuning, batch size=32k2023.07 | 84.5 | — | |
| DenseSparsity (%)=0, Backbone=BERT-base2023.05 | 84.5 | — | |
| Annealing-KD#aug=-2023.05 | 84.5 | — | |
| Baselinemode=fine-tuning2023.07 | 84.3 | — | |
| WOAens#aug=5k, ensemble=true2023.05 | 84.3 | — | |
| DMU full#aug=-2023.05 | 84.2 | — | |
| KD#aug=-2023.05 | 84.1 | — | |
| DistilRoBERTa2023.05 | 83.8 | — | |
| PDPSparsity (%)=90, Backbone=BERT-base2023.05 | 83.1 | — | |
| PDPSparsity (%)=94, Backbone=BERT-base2023.05 | 82 | — | |
| WOA**ens#aug=5k, ensemble=true, distillation_step_on_augmented_data_only=true2023.05 | 81.6 | — | |
| POFASparsity (%)=90, Backbone=BERT-base2023.05 | 81.5 | — | |
| MVPSparsity (%)=90, Backbone=BERT-base2023.05 | 81.2 | — | |
| MVPSparsity (%)=94, Backbone=BERT-base2023.05 | 80.7 | — | |
| LIRExBackbone=RoBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 79.85 | -0.27 | |
| DPS MixTraining Dataset=MNLI2022.11 | 79.16 | — | |
| CHILD-TUNINGDTraining Dataset=MNLI2022.11 | 79.13 | — | |
| DPS DenseTraining Dataset=MNLI2022.11 | 79.03 | — | |
| vanillaTraining Dataset=MNLI2022.11 | 78.75 | — | |
| OptGSparsity (%)=90, Backbone=BERT-base2023.05 | 78.5 | — | |
| NILEBackbone=RoBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 77.07 | -2.22 | |
| OptGSparsity (%)=94, Backbone=BERT-base2023.05 | 76.9 | — | |
| DTAensBackbone=BERT, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 76.45 | 1.48 | |
| STRSparsity (%)=90, Backbone=BERT-base2023.05 | 75.8 | — | |
| Prompt-based FT (hard) + PCPBackbone=RoBERTa-LARGE, Supervision Setting=Semi-supervised (16-shot), Evaluation Protocol=Prompt-based Fine-tuning (hard), Continued Pre-training=PCP2023.05 | 75.6 | — | |
| KDBackbone=BERT, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 75.5 | 0.53 | |
| BERT teacherRole=Teacher, Architecture=BERT2023.05 | 74.97 | — | |
| GMPSparsity (%)=94, Backbone=BERT-base2023.05 | 74.8 | — | |
| STRSparsity (%)=94, Backbone=BERT-base2023.05 | 74.4 | — | |
| Product of ExpertsBackbone=BERT, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 73.61 | -0.79 | |
| Debiased Focal LossBackbone=BERT, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 73.58 | -0.82 | |
| Rationale supervisionBackbone=BERT, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 73.19 | 0.91 | |
| Meta-training with Demonstration Retrievalmeta-trained=true2023.06 | 72.9 | — | |
| OMModel Type=Sequence Encoder2023.11 | 72.5 | — | |
| CRvNNModel Type=Sequence Encoder2023.11 | 72.2 | — | |
| EBT-GRCModel Type=Sequence Encoder2023.11 | 72.1 | — | |
| RIR-EBT-GRCModel Type=Sequence Encoder2023.11 | 71.8 | — | |
| BT-GRC OSModel Type=Sequence Encoder2023.11 | 71.7 | — | |
| RIR-GRCModel Type=Sequence Encoder2023.11 | 71.5 | — | |
| iPET2023.06 | 71.2 | — | |
| BBT-GRCModel Type=Sequence Encoder2023.11 | 71.1 | — | |
| LM-BFF2023.06 | 70.7 | — | |
| RAGmeta-trained=true2023.06 | 70 | — | |
| CoTAMEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 69.7 | — | |
| DPS DenseTraining Dataset=SNLI2022.11 | 69.08 | — | |
| FlipDA++Evaluation Protocol=ICL, Number of shots (K)=32023.07 | 68.8 | — | |
| CHILD-TUNINGDTraining Dataset=SNLI2022.11 | 68.75 | — | |
| Extra AnnotationEvaluation Protocol=ICL, Number of shots (K)=Increased2023.07 | 68.7 | — | |
| DPS MixTraining Dataset=SNLI2022.11 | 68.47 | — | |
| COTDAEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 68.2 | — | |
| BaseEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 68.1 | — | |
| vanillaTraining Dataset=SNLI2022.11 | 68.07 | — | |
| No ExampleEvaluation Protocol=ICL, Number of shots (K)=02023.07 | 67.5 | — | |
| LLM Pseudo LabelEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 66.9 | — | |
| OPT-IML 175BEvaluation protocol=few-shot, k=52022.12 | 64.4 | — | |
| RAGmeta-trained=false2023.06 | 62.4 | — | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 61.1 | — | |
| FLAN 137BEvaluation protocol=few-shot, k=variable, Training strategy=leave-one-category-out2022.12 | 60.8 | — | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 59.2 | — | |
| DTARole=Student, Architecture=TinyBERT2023.05 | 57.17 | — | |
| DTA with DMURole=Student, Architecture=TinyBERT2023.05 | 56.54 | — | |
| KD + SmoothingRole=Student, Architecture=TinyBERT2023.05 | 55.83 | — | |
| KD (standard distillation)Role=Student, Architecture=TinyBERT2023.05 | 55.82 | — | |
| DecTn (Shots)=162022.12 | 55.3 | — | |
| TinyBERT baselineRole=Student, Architecture=TinyBERT2023.05 | 54.24 | — | |
| Ensemble adversariesBackbone=LSTM, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 54.18 | 0.8 | |
| DMURole=Student, Architecture=TinyBERT2023.05 | 54.08 | — | |
| CoTAMEvaluation Protocol=Fine-tuning, Number of shots (K)=10, Backbone=RoBERTa-Large2023.07 | 54.07 | — | |
| OPT-IML 30BEvaluation protocol=few-shot, k=52022.12 | 53.6 | — | |
| JTTRole=Student, Architecture=TinyBERT2023.05 | 52 | — | |
| FlipDA++Evaluation Protocol=Fine-tuning, Number of shots (K)=10, Backbone=RoBERTa-Large2023.07 | 51.52 | — | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 51.1 | — | |
| Promptn (Shots)=02022.12 | 50.8 | — | |
| FewshotQAmeta-trained=true2023.06 | 50.1 | — | |
| FewshotQAmeta-trained=false2023.06 | 47.9 | — | |
| Hyp-only adversaryBackbone=LSTM, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 47.24 | 1.38 | |
| RoBERTa2023.06 | 45.8 | — | |
| Baseline w/ labelled aug. dataRole=Student, Architecture=TinyBERT2023.05 | 45.55 | — | |
| Negative samplingBackbone=LSTM, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 43.76 | -2.1 | |
| LaMDA-PT 137BEvaluation protocol=few-shot, k=variable2022.12 | 43.7 | — |