Natural Language Inference on WNLI
97.9AccuracySOTA
Evaluation Results
| Method | Links | |
|---|---|---|
| SOTA2023.05 | 97.9 | |
| GPT-4Prompting=Standard GPT-style, Shot=Zero-shot2023.05 | 91.6 | |
| MUPPET2022.12 | 91.1 | |
| Knowledge Hunter2019.05 | 90.1 | |
| ChatGPTPrompting=Standard GPT-style, Shot=Zero-shot2023.05 | 81.7 | |
| BERT_WIKI_WSCRPre-training=MaskedWiki, Fine-tuned on WSCR=true2019.05 | 74.7 | |
| Training-free Weight Recycling (L^T_KD)Distillation Data=Unlabeled Task Data, Source Dataset=QNLI2023.05 | 72.2 | |
| BERT_WSCRFine-tuned on WSCR=true2019.05 | 71.9 | |
| BERT_WIKIPre-training=MaskedWiki, Fine-tuned on WSCR=false2019.05 | 71.2 | |
| BERT-base_WSCRFine-tuned on WSCR=true2019.05 | 70.5 | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 70 | |
| BERTFine-tuned on WSCR=false2019.05 | 65.8 | |
| LaMDA-PT 137BEvaluation protocol=few-shot, k=variable2022.12 | 64.8 | |
| BERT-baseFine-tuned on WSCR=false2019.05 | 63 | |
| OPT-IML 175BEvaluation protocol=few-shot, k=52022.12 | 62.7 | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 61 | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 58.5 | |
| Training-free Weight Recycling (L^W_KD)Distillation Data=Wikipedia, Source Dataset=QNLI2023.05 | 58.3 | |
| RPJ-INCITE-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 57.8 | |
| OPT-IML 30BEvaluation protocol=few-shot, k=52022.12 | 57.7 | |
| LaMDA-PT 137BEvaluation protocol=0-shot2022.12 | 56.3 | |
| OLMo-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 56.3 | |
| Demonstration learningMode=Zero-shot with context2023.05 | 55.6 | |
| Full dataset tuningTraining=Full dataset (FD)2023.05 | 55.6 | |
| FLAN 137BEvaluation protocol=few-shot, k=variable, Training strategy=leave-one-category-out2022.12 | 55.4 | |
| OPT 175BEvaluation protocol=0-shot2022.12 | 55.4 | |
| Finetune2022.12 | 55.21 | |
| ColD-Fusion2022.12 | 54.93 | |
| LLaMA-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 52.1 | |
| Multitask2022.12 | 51.55 | |
| OPT 30BEvaluation protocol=few-shot, k=52022.12 | 50.6 | |
| OPT 30BEvaluation protocol=0-shot2022.12 | 50.3 | |
| 32-shot tuningTraining=32-shot (FS)2023.05 | 50 | |
| RWKV-4-Raven-14BPrompting=Adapted (re-ordered), Shot=Zero-shot2023.05 | 49.3 | |
| RWKV-4-Raven-14BPrompting=Standard GPT-style, Shot=Zero-shot2023.05 | 47.9 | |
| Falcon-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 47.9 | |
| MPT-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 47.9 | |
| OPT 175BEvaluation protocol=few-shot, k=52022.12 | 47.7 | |
| LLaMA2-7BEvaluation protocol=Zero-shot, Parameters=7B2024.02 | 45.1 | |
| Pythia-6.9BEvaluation protocol=Zero-shot, Parameters=6.9B2024.02 | 38 |