Natural Language Inference on ANLI Round 1
77AccuracyFLAN-T5
Evaluation Results
| Method | Links | |
|---|---|---|
| FLAN-T5Model Variant=xlarge, Model Parameters=3B2023.07 | 77 | |
| ALIGNModel Variant=large, Model Parameters=355M2023.07 | 75.8 | |
| Traditional MTL2024.05 | 70.5 | |
| Individual2024.05 | 70.2 | |
| FLAN-T5Model Variant=large, Model Parameters=780M2023.07 | 68.1 | |
| Ties-MergingValidation=true2024.05 | 66.9 | |
| EMR-MERGINGValidation=false2024.05 | 65.7 | |
| ALIGNModel Variant=base, Model Parameters=125M2023.07 | 65.3 | |
| Task ArithmeticValidation=true2024.05 | 60.8 | |
| Task ArithmeticValidation=false2024.05 | 59.8 | |
| Ties-MergingValidation=false2024.05 | 58.1 | |
| OPT-IML 175BEvaluation protocol=few-shot, k=52022.12 | 48 | |
| FLAN 137BEvaluation protocol=0-shot, Training strategy=leave-one-category-out2022.12 | 47.7 | |
| OPT-IML 175BEvaluation protocol=0-shot2022.12 | 46.1 | |
| Fisher MergingValidation=true2024.05 | 45.9 | |
| FLAN 137BEvaluation protocol=few-shot, k=variable, Training strategy=leave-one-category-out2022.12 | 44.2 | |
| T0-11BParameters=11B2023.02 | 43.56 | |
| RegMeanValidation=true2024.05 | 43.3 | |
| Weight AveragingValidation=false2024.05 | 43.3 | |
| T5(3B) + PE w/ ROE (ORC.)Backbone=T5 (3B), Expert type=Oracle Retrieval-of-Expert, Additional Parameters=100M2023.02 | 40.02 | |
| LaMDA-PT 137BEvaluation protocol=0-shot2022.12 | 39.6 | |
| LaMDA-PT 137BEvaluation protocol=few-shot, k=variable2022.12 | 39 | |
| OPT-IML 30BEvaluation protocol=0-shot2022.12 | 38.5 | |
| OPT-IML 30BEvaluation protocol=few-shot, k=52022.12 | 36.5 | |
| GPT-3Model Variant=DaVinci, Zero-shot=true2022.04 | 36.3 | |
| T5(3B) + Cos PEBackbone=T5 (3B), Expert type=Prompt Expert trained on COSMOS-QA, Additional Parameters=100M2023.02 | 36.21 | |
| T5(3B) + PE w/ ROEBackbone=T5 (3B), Expert type=Retrieval-of-Expert, Additional Parameters=100M2023.02 | 35.49 | |
| T0-3BParameters=3B2023.02 | 35.1 | |
| GPT-3 (175B)Parameters=175B2023.02 | 34.6 | |
| FairSeqNumber of Parameters=13B, Zero-Shot=true2022.04 | 34 | |
| GPT-NeoXModel Size=20B, Zero-shot=true2022.04 | 34 | |
| OPT 175BEvaluation protocol=few-shot, k=52022.12 | 34 | |
| FairSeqNumber of Parameters=6.7B, Zero-Shot=true2022.04 | 33.8 | |
| FairSeqNumber of Parameters=355M, Evaluation Protocol=5-shot2022.04 | 33.6 | |
| FairSeqNumber of Parameters=2.7B, Evaluation Protocol=5-shot2022.04 | 33.6 | |
| BLOOM-176BEvaluation Protocol=1-shot2023.03 | 33.6 | |
| FairSeqNumber of Parameters=13B, Evaluation Protocol=5-shot2022.04 | 33.5 | |
| GPT-3Model Variant=Ada, Zero-shot=true2022.04 | 33.4 | |
| OPT 30BEvaluation protocol=0-shot2022.12 | 33.3 | |
| OPT 30BEvaluation protocol=few-shot, k=52022.12 | 33.3 | |
| OPT 175BEvaluation protocol=0-shot2022.12 | 33.3 | |
| FairSeqNumber of Parameters=125M, Evaluation Protocol=5-shot2022.04 | 33.2 | |
| FairSeqNumber of Parameters=1.3B, Zero-Shot=true2022.04 | 33.1 | |
| OPT-66BEvaluation Protocol=1-shot2023.03 | 33.1 | |
| BloombergGPTEvaluation Protocol=1-shot2023.03 | 32.9 | |
| FairSeqNumber of Parameters=1.3B, Evaluation Protocol=5-shot2022.04 | 32.7 | |
| GPT-3Model Variant=Babbage, Zero-shot=true2022.04 | 32.6 | |
| GPT-NeoXEvaluation Protocol=1-shot2023.03 | 32.6 | |
| GPT-3Model Variant=Curie, Zero-shot=true2022.04 | 32.5 | |
| GPT-JModel Size=6B, Zero-shot=true2022.04 | 32.4 | |
| FairSeqNumber of Parameters=355M, Zero-Shot=true2022.04 | 32.2 | |
| GPT-JParameters=6B, Evaluation=Five-shot2022.04 | 32.2 | |
| GPT-3Evaluation Protocol=1-shot2023.03 | 32 | |
| FairSeqNumber of Parameters=2.7B, Zero-Shot=true2022.04 | 31.8 | |
| FairSeqNumber of Parameters=125M, Zero-Shot=true2022.04 | 31.6 | |
| GPT-NeoXParameters=20B, Evaluation=Five-shot2022.04 | 31.2 | |
| FairSeqNumber of Parameters=6.7B, Evaluation Protocol=5-shot2022.04 | 30.5 |