Natural Language Inference on ANLI (test)
92.2Overall ScoreTEAM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TEAMBackbone=DeBERTa Large2022.10 | 92.2 | — | — | — | |
| ScoreBackbone=DeBERTa Large2022.10 | 89.74 | — | — | — | |
| UNICORN 11BModel=UNICORN 11B2022.10 | 87.3 | — | — | — | |
| TEAMBackbone=RoBERTa Large2022.10 | 87.04 | — | — | — | |
| ScoreBackbone=RoBERTa Large2022.10 | 83.91 | — | — | — | |
| InfoBERTTraining=Adversarial Training, Model=RoBERTa2020.10 | 58.3 | 75.5 | 51.4 | 49.8 | |
| InfoBERTTraining=Standard Training, Model=RoBERTa2020.10 | 57.3 | 73.9 | 50.8 | 48.8 | |
| SMART_RoBERTa-LARGETraining Data=MNLI + SNLI + ANLI + FEVER2019.11 | 57.1 | 72.4 | 49.8 | 50.3 | |
| SMARTTraining=Adversarial Training, Model=RoBERTa2020.10 | 57.1 | 72.4 | 49.8 | 50.3 | |
| ALUMTraining=Adversarial Training, Model=RoBERTa2020.10 | 57 | 72.3 | 52.1 | 48.4 | |
| SMART_RoBERTa-LARGETraining Data=ANLI2019.11 | 56.9 | 72.4 | 50.3 | 49.5 | |
| FreeLBTraining=Adversarial Training, Model=RoBERTa2020.10 | 56.2 | 73.3 | 50.5 | 46.8 | |
| XLNet_LARGETraining Data=MNLI + SNLI + ANLI + FEVER2019.11 | 55.1 | 67.6 | 50.7 | 48.3 | |
| ICLmode=Zero-shot transfer, context_examples=Sampled from MNLI2023.05 | 54.69 | 59.5 | 52.4 | 52.58 | |
| RoBERTa_LARGETraining Data=MNLI + SNLI + ANLI + FEVER2019.11 | 53.7 | 73.8 | 48.9 | 44.4 | |
| VanillaTraining=Standard Training, Model=RoBERTa2020.10 | 53.7 | 73.8 | 48.9 | 44.4 | |
| RoBERTa_LARGETraining Data=ANLI2019.11 | 51.9 | 71.3 | 43.3 | 43 | |
| InfoBERTTraining=Adversarial Training, Model=BERT2020.10 | 51.2 | 63.3 | 48.7 | 43.2 | |
| InfoBERTTraining=Standard Training, Model=BERT2020.10 | 50.2 | 60 | 46.9 | 44.8 | |
| FreeLBTraining=Adversarial Training, Model=BERT2020.10 | 50.2 | 60.3 | 46.8 | 44.8 | |
| ALUMTraining=Adversarial Training, Model=BERT2020.10 | 50.1 | 61.3 | 45.9 | 44.3 | |
| BERT_LARGETraining Data=MNLI + SNLI + ANLI + FEVER2019.11 | 49.3 | 57.4 | 48.3 | 43.5 | |
| VanillaTraining=Standard Training, Model=BERT2020.10 | 49.3 | 57.4 | 48.3 | 43.5 | |
| SuperICLmode=Zero-shot transfer, context_examples=Sampled from MNLI2023.05 | 47.44 | 56.1 | 42.7 | 44.17 | |
| MI-based distillationBackbone=T5-small2024.03 | 43.7 | — | — | — | |
| DSSBackbone=T5-small, Evaluation Protocol=Distilling Step-by-Step2024.03 | 42.9 | — | — | — | |
| FinetuningBackbone=T5-small, Evaluation Protocol=Standard Label Supervision2024.03 | 42 | — | — | — | |
| RoBERTa-Largetraining=Fine-tuned on MNLI2023.05 | 30.78 | 41.6 | 27.4 | 24.58 |