Natural Language Inference on ANLI Round 3
67.9AccuracyLMSI
Evaluation Results
| Method | Links | |
|---|---|---|
| LMSIPrompting Method=Self-Consistency2022.10 | 67.9 | |
| LMSIPrompting Method=CoT-Prompting2022.10 | 67.3 | |
| LMSIPrompting Method=Standard-Prompting2022.10 | 66.9 | |
| Wang et al. (2022a)description=Previous SOTA2022.10 | 66 | |
| Self-ConsistencyLMSI=false2022.10 | 63.4 | |
| CoT-PromptingLMSI=false2022.10 | 60.6 | |
| FLAN-T5Model Variant=xlarge, Model Parameters=3B2023.07 | 56.6 | |
| Standard-PromptingLMSI=false2022.10 | 55.8 | |
| Individual2024.05 | 53 | |
| ALIGNModel Variant=large, Model Parameters=355M2023.07 | 52.3 | |
| Ties-MergingValidation=true2024.05 | 51.1 | |
| EMR-MERGINGValidation=false2024.05 | 50.8 | |
| Task ArithmeticValidation=true2024.05 | 50 | |
| FLAN-T5Model Variant=large, Model Parameters=780M2023.07 | 49.8 | |
| Task ArithmeticValidation=false2024.05 | 48.2 | |
| Traditional MTL2024.05 | 47.7 | |
| Ties-MergingValidation=false2024.05 | 47.4 | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 47 | |
| ALIGNModel Variant=base, Model Parameters=125M2023.07 | 45.5 | |
| OPT-IML 175BEvaluation protocol=few-shot, k=52022.12 | 44.1 | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 43.8 | |
| FLAN 137BEvaluation protocol=few-shot, k=variable, Training strategy=leave-one-category-out2022.12 | 42.8 | |
| Fisher MergingValidation=true2024.05 | 42.2 | |
| T5(3B) + PE w/ ROE (ORC.)Backbone=T5 (3B), Expert type=Oracle Retrieval-of-Expert, Additional Parameters=100M2023.02 | 42.07 | |
| T0-11BParameters=11B2023.02 | 41.26 | |
| LaMDA-PT 137BEvaluation protocol=few-shot, k=variable2022.12 | 40.7 | |
| RegMeanValidation=true2024.05 | 40.2 | |
| Weight AveragingValidation=false2024.05 | 40.2 | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 39.6 | |
| LaMDA-PT 137BEvaluation protocol=0-shot2022.12 | 39.3 | |
| OPT-IML 30BEvaluation protocol=few-shot, k=52022.12 | 38.3 | |
| BloombergGPTEvaluation Protocol=1-shot2023.03 | 37.33 | |
| FairSeqNumber of Parameters=1.3B, Evaluation Protocol=5-shot2022.04 | 37 | |
| GPT-3Model Variant=DaVinci, Zero-shot=true2022.04 | 36.9 | |
| FairSeqNumber of Parameters=6.7B, Evaluation Protocol=5-shot2022.04 | 36.7 | |
| T5(3B) + Cos PEBackbone=T5 (3B), Expert type=Prompt Expert trained on COSMOS-QA, Additional Parameters=100M2023.02 | 36.38 | |
| GPT-NeoXEvaluation Protocol=1-shot2023.03 | 36.17 | |
| FairSeqNumber of Parameters=125M, Evaluation Protocol=5-shot2022.04 | 35.9 | |
| FairSeqNumber of Parameters=13B, Evaluation Protocol=5-shot2022.04 | 35.7 | |
| GPT-JModel Size=6B, Zero-shot=true2022.04 | 35.5 | |
| GPT-NeoXModel Size=20B, Zero-shot=true2022.04 | 35.4 | |
| GPT-3Model Variant=Ada, Zero-shot=true2022.04 | 35.4 | |
| GPT-3Model Variant=Curie, Zero-shot=true2022.04 | 35.3 | |
| BLOOM-176BEvaluation Protocol=1-shot2023.03 | 35.17 | |
| GPT-3Evaluation Protocol=1-shot2023.03 | 35.1 | |
| OPT-66BEvaluation Protocol=1-shot2023.03 | 34.92 | |
| FairSeqNumber of Parameters=13B, Zero-Shot=true2022.04 | 34.7 | |
| FairSeqNumber of Parameters=355M, Evaluation Protocol=5-shot2022.04 | 34.7 | |
| GPT-JParameters=6B, Evaluation=Five-shot2022.04 | 34.6 | |
| OPT 175BEvaluation protocol=few-shot, k=52022.12 | 34.6 | |
| GPT-3 (175B)Parameters=175B2023.02 | 34.5 | |
| GPT-NeoXParameters=20B, Evaluation=Five-shot2022.04 | 34.2 | |
| FairSeqNumber of Parameters=2.7B, Zero-Shot=true2022.04 | 34 | |
| GPT-3Model Variant=Babbage, Zero-shot=true2022.04 | 34 | |
| OPT 30BEvaluation protocol=0-shot2022.12 | 33.5 | |
| OPT 30BEvaluation protocol=few-shot, k=52022.12 | 33.5 | |
| OPT 175BEvaluation protocol=0-shot2022.12 | 33.5 | |
| T0-3BParameters=3B2023.02 | 33.32 | |
| FairSeqNumber of Parameters=1.3B, Zero-Shot=true2022.04 | 33.3 | |
| FairSeqNumber of Parameters=6.7B, Zero-Shot=true2022.04 | 33.3 | |
| FairSeqNumber of Parameters=125M, Zero-Shot=true2022.04 | 33 | |
| FairSeqNumber of Parameters=2.7B, Evaluation Protocol=5-shot2022.04 | 32.6 | |
| FairSeqNumber of Parameters=355M, Zero-Shot=true2022.04 | 32.3 | |
| T5(3B) + PE w/ ROEBackbone=T5 (3B), Expert type=Retrieval-of-Expert, Additional Parameters=100M2023.02 | 31.22 |