Natural Language Inference on ANLI R2 (test)
33.1AccuracyChaosNLI HJD
Evaluation Results
| Method | Links | |
|---|---|---|
| ChaosNLI HJDEvaluation Model=RoBERTa2024.12 | 33.1 | |
| VariErr distributionEvaluation Model=RoBERTa2024.12 | 31.1 | |
| Mixtral (MJD) + human explanationsEvaluation Model=RoBERTa, Explanations=Human2024.12 | 29.2 | |
| ChaosNLI HJDEvaluation Model=BERT2024.12 | 28.9 | |
| VariErr Label-GuidedEvaluation Model=RoBERTa, Explanations=Model-generated2024.12 | 28.9 | |
| VariErr Label-GuidedEvaluation Model=BERT, Explanations=Model-generated2024.12 | 28.7 | |
| Mixtral (MJD) + human explanationsEvaluation Model=BERT, Explanations=Human2024.12 | 28 | |
| MNLI Label-GuidedEvaluation Model=RoBERTa, Explanations=Model-generated2024.12 | 28 | |
| MNLI distributionEvaluation Model=RoBERTa2024.12 | 27.5 | |
| MNLI Label-GuidedEvaluation Model=BERT, Explanations=Model-generated2024.12 | 27.3 | |
| MNLI-FT-LMEvaluation Model=BERT2024.12 | 26.9 | |
| MNLI-FT-LMEvaluation Model=RoBERTa2024.12 | 26.2 | |
| MNLI distributionEvaluation Model=BERT2024.12 | 26 | |
| VariErr distributionEvaluation Model=BERT2024.12 | 25.9 | |
| Label-FreeEvaluation Model=BERT, Explanations=Model-generated2024.12 | 25.5 | |
| Mixtral (MJD)Evaluation Model=BERT, Explanations=None2024.12 | 25.2 | |
| Label-FreeEvaluation Model=RoBERTa, Explanations=Model-generated2024.12 | 24.8 | |
| Mixtral (MJD)Evaluation Model=RoBERTa, Explanations=None2024.12 | 24 | |
| Out-of-the-box LMEvaluation Model=BERT2024.12 | 17.6 | |
| Out-of-the-box LMEvaluation Model=RoBERTa2024.12 | 16.7 |