Natural Language Inference on ANLI Round 2
66.5AccuracyLMSI
Evaluation Results
| Method | Links | |
|---|---|---|
| LMSIPrompting Method=Self-Consistency2022.10 | 66.5 | |
| LMSIPrompting Method=CoT-Prompting2022.10 | 65.3 | |
| Wang et al. (2022a)description=Previous SOTA2022.10 | 64.9 | |
| LMSIPrompting Method=Standard-Prompting2022.10 | 64.8 | |
| Self-ConsistencyLMSI=false2022.10 | 64.5 | |
| FLAN-T5Model Variant=xlarge, Model Parameters=3B2023.07 | 60.6 | |
| CoT-PromptingLMSI=false2022.10 | 58.9 | |
| Standard-PromptingLMSI=false2022.10 | 55.8 | |
| ALIGNModel Variant=large, Model Parameters=355M2023.07 | 52.4 | |
| Ties-MergingValidation=true2024.05 | 51.3 | |
| Traditional MTL2024.05 | 49.8 | |
| Task ArithmeticValidation=true2024.05 | 49.4 | |
| ALIGNModel Variant=base, Model Parameters=125M2023.07 | 48.7 | |
| FLAN-T5Model Variant=large, Model Parameters=780M2023.07 | 48.7 | |
| Task ArithmeticValidation=false2024.05 | 47.5 | |
| Individual2024.05 | 46.5 | |
| Ties-MergingValidation=false2024.05 | 46.5 | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 43.9 | |
| OPT-IML 175BEvaluation protocol=few-shot, k=52022.12 | 43.8 | |
| EMR-MERGINGValidation=false2024.05 | 43.8 | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 43.5 | |
| FLAN 137BEvaluation protocol=few-shot, k=variable, Training strategy=leave-one-category-out2022.12 | 41.6 | |
| Fisher MergingValidation=true2024.05 | 41 | |
| T5(3B) + PE w/ ROE (ORC.)Backbone=T5 (3B), Expert type=Oracle Retrieval-of-Expert, Additional Parameters=100M2023.02 | 40.11 | |
| LaMDA-PT 137BEvaluation protocol=0-shot2022.12 | 39.9 | |
| RegMeanValidation=true2024.05 | 39.2 | |
| Weight AveragingValidation=false2024.05 | 39.2 | |
| T0-11BParameters=11B2023.02 | 38.68 | |
| GPT-3Model Variant=DaVinci, Zero-shot=true2022.04 | 37.5 | |
| LaMDA-PT 137BEvaluation protocol=few-shot, k=variable2022.12 | 37.5 | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 37.5 | |
| OPT-IML 30BEvaluation protocol=few-shot, k=52022.12 | 37 | |
| T5(3B) + Cos PEBackbone=T5 (3B), Expert type=Prompt Expert trained on COSMOS-QA, Additional Parameters=100M2023.02 | 36.11 | |
| GPT-3 (175B)Parameters=175B2023.02 | 35.4 | |
| FairSeqNumber of Parameters=355M, Evaluation Protocol=5-shot2022.04 | 35 | |
| OPT 175BEvaluation protocol=few-shot, k=52022.12 | 35 | |
| FairSeqNumber of Parameters=1.3B, Evaluation Protocol=5-shot2022.04 | 34.7 | |
| T5(3B) + PE w/ ROEBackbone=T5 (3B), Expert type=Retrieval-of-Expert, Additional Parameters=100M2023.02 | 34.64 | |
| FairSeqNumber of Parameters=125M, Evaluation Protocol=5-shot2022.04 | 34.5 | |
| BloombergGPTEvaluation Protocol=1-shot2023.03 | 34.4 | |
| GPT-NeoXModel Size=20B, Zero-shot=true2022.04 | 34.3 | |
| GPT-3Model Variant=Ada, Zero-shot=true2022.04 | 34.2 | |
| OPT-66BEvaluation Protocol=1-shot2023.03 | 34.2 | |
| FairSeqNumber of Parameters=6.7B, Evaluation Protocol=5-shot2022.04 | 34 | |
| GPT-JModel Size=6B, Zero-shot=true2022.04 | 34 | |
| FairSeqNumber of Parameters=2.7B, Zero-Shot=true2022.04 | 33.9 | |
| GPT-3Evaluation Protocol=1-shot2023.03 | 33.9 | |
| FairSeqNumber of Parameters=13B, Evaluation Protocol=5-shot2022.04 | 33.8 | |
| GPT-3Model Variant=Curie, Zero-shot=true2022.04 | 33.8 | |
| GPT-NeoXEvaluation Protocol=1-shot2023.03 | 33.8 | |
| BLOOM-176BEvaluation Protocol=1-shot2023.03 | 33.8 | |
| FairSeqNumber of Parameters=125M, Zero-Shot=true2022.04 | 33.6 | |
| OPT 30BEvaluation protocol=few-shot, k=52022.12 | 33.6 | |
| FairSeqNumber of Parameters=1.3B, Zero-Shot=true2022.04 | 33.4 | |
| FairSeqNumber of Parameters=2.7B, Evaluation Protocol=5-shot2022.04 | 33.3 | |
| OPT 30BEvaluation protocol=0-shot2022.12 | 33.3 | |
| OPT 175BEvaluation protocol=0-shot2022.12 | 33.3 | |
| T0-3BParameters=3B2023.02 | 33.27 | |
| GPT-JParameters=6B, Evaluation=Five-shot2022.04 | 33.1 | |
| FairSeqNumber of Parameters=13B, Zero-Shot=true2022.04 | 33 | |
| GPT-NeoXParameters=20B, Evaluation=Five-shot2022.04 | 32.9 | |
| FairSeqNumber of Parameters=6.7B, Zero-Shot=true2022.04 | 32.2 | |
| FairSeqNumber of Parameters=355M, Zero-Shot=true2022.04 | 31.2 | |
| GPT-3Model Variant=Babbage, Zero-shot=true2022.04 | 30.8 |