Natural Language Inference on MNLI (mismatched)
91AccuracyMETALM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| METALMtuning_mode=Single-task finetuning, parameters_updated=non-causal encoder and connector, frozen_components=language model2022.06 | 91 | — | |
| FLAN-T5Model Variant=xlarge, Model Parameters=3B2023.07 | 90.4 | — | |
| ALIGNModel Variant=large, Model Parameters=355M2023.07 | 90.3 | — | |
| RoBERTatuning_mode=Single-task finetuning, model_size=Large2022.06 | 90.2 | — | |
| FLAN-T5Model Variant=large, Model Parameters=780M2023.07 | 88.7 | — | |
| GPTtuning_mode=Single-task finetuning, parameters_updated=all2022.06 | 87.6 | — | |
| ALIGNModel Variant=base, Model Parameters=125M2023.07 | 87.5 | — | |
| DTAensBackbone=DeBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 86.18 | 1.4 | |
| KDBackbone=DeBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 85.24 | 0.46 | |
| DenseSparsity (%)=0, Backbone=BERT-base2023.05 | 84.9 | — | |
| PDPSparsity (%)=90, Backbone=BERT-base2023.05 | 83 | — | |
| POFASparsity (%)=90, Backbone=BERT-base2023.05 | 82.4 | — | |
| PDPSparsity (%)=94, Backbone=BERT-base2023.05 | 82.4 | — | |
| MVPSparsity (%)=90, Backbone=BERT-base2023.05 | 81.8 | — | |
| MVPSparsity (%)=94, Backbone=BERT-base2023.05 | 81.2 | — | |
| LIRExBackbone=RoBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 79.79 | 0.06 | |
| OptGSparsity (%)=90, Backbone=BERT-base2023.05 | 78.3 | — | |
| NILEBackbone=RoBERTa, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 77.22 | -2.07 | |
| OptGSparsity (%)=94, Backbone=BERT-base2023.05 | 76.5 | — | |
| DTAensBackbone=BERT, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 76.42 | 1.41 | |
| STRSparsity (%)=90, Backbone=BERT-base2023.05 | 76.3 | — | |
| GMPSparsity (%)=94, Backbone=BERT-base2023.05 | 75.6 | — | |
| KDBackbone=BERT, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 75.42 | 0.41 | |
| STRSparsity (%)=94, Backbone=BERT-base2023.05 | 74.1 | — | |
| Debiased Focal LossBackbone=BERT, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 74 | 0.02 | |
| Product of ExpertsBackbone=BERT, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 73.49 | -0.49 | |
| Rationale supervisionBackbone=BERT, Hyper-parameter tuning on MNLI=false, Zero-shot=true2023.05 | 73.36 | 0.84 | |
| OMModel Type=Sequence Encoder2023.11 | 73.2 | — | |
| CRvNNModel Type=Sequence Encoder2023.11 | 72.6 | — | |
| RIR-EBT-GRCModel Type=Sequence Encoder2023.11 | 72.3 | — | |
| EBT-GRCModel Type=Sequence Encoder2023.11 | 72.1 | — | |
| LM-BFF2023.06 | 72 | — | |
| BT-GRC OSModel Type=Sequence Encoder2023.11 | 71.9 | — | |
| iPET2023.06 | 71.8 | — | |
| RIR-GRCModel Type=Sequence Encoder2023.11 | 71.6 | — | |
| BBT-GRCModel Type=Sequence Encoder2023.11 | 71.4 | — | |
| No ExampleEvaluation Protocol=ICL, Number of shots (K)=02023.07 | 69.7 | — | |
| Meta-training with Demonstration Retrievalmeta-trained=true2023.06 | 69.6 | — | |
| CoTAMEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 69.2 | — | |
| RAGmeta-trained=true2023.06 | 69.1 | — | |
| LLM Pseudo LabelEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 69 | — | |
| FlipDA++Evaluation Protocol=ICL, Number of shots (K)=32023.07 | 68.9 | — | |
| Extra AnnotationEvaluation Protocol=ICL, Number of shots (K)=Increased2023.07 | 68.6 | — | |
| COTDAEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 68.5 | — | |
| BaseEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 68.1 | — | |
| RIR-EBT-GRCModel Type=Sequence Encoder2023.11 | 67.3 | — | |
| BT-GRC OSModel Type=Sequence Encoder2023.11 | 66.7 | — | |
| EBT-GRCModel Type=Sequence Encoder2023.11 | 66.4 | — | |
| CRvNNModel Type=Sequence Encoder2023.11 | 63.3 | — | |
| RAGmeta-trained=false2023.06 | 61.8 | — | |
| BBT-GRCModel Type=Sequence Encoder2023.11 | 61.8 | — | |
| OMModel Type=Sequence Encoder2023.11 | 57 | — | |
| DecTn (Shots)=162022.12 | 56.8 | — | |
| CoTAMEvaluation Protocol=Fine-tuning, Number of shots (K)=10, Backbone=RoBERTa-Large2023.07 | 56.16 | — | |
| RIR-GRCModel Type=Sequence Encoder2023.11 | 55.8 | — | |
| FlipDA++Evaluation Protocol=Fine-tuning, Number of shots (K)=10, Backbone=RoBERTa-Large2023.07 | 53.56 | — | |
| Ensemble adversariesBackbone=LSTM, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 52.81 | -0.1 | |
| Promptn (Shots)=02022.12 | 51.7 | — | |
| FewshotQAmeta-trained=true2023.06 | 50.6 | — | |
| Hyp-only adversaryBackbone=LSTM, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 49.24 | 1.67 | |
| RoBERTa2023.06 | 47.8 | — | |
| FewshotQAmeta-trained=false2023.06 | 46.1 | — | |
| Extra AnnotationEvaluation Protocol=Fine-tuning, Number of shots (K)=Increased, Backbone=RoBERTa-Large2023.07 | 44.03 | — | |
| Negative samplingBackbone=LSTM, Hyper-parameter tuning on MNLI=true, Zero-shot=true2023.05 | 43.66 | -3.91 | |
| LLM Pseudo LabelEvaluation Protocol=Fine-tuning, Number of shots (K)=10, Backbone=RoBERTa-Large2023.07 | 42.92 | — | |
| BaseEvaluation Protocol=Fine-tuning, Number of shots (K)=10, Backbone=RoBERTa-Large2023.07 | 38.75 | — | |
| COTDAEvaluation Protocol=Fine-tuning, Number of shots (K)=10, Backbone=RoBERTa-Large2023.07 | 36.28 | — | |
| Majority2023.06 | 33.3 | — |