Natural Language Inference on E-SNLI
91.31AccuracyColD-Fusion
Evaluation Results
| Method | Links | |
|---|---|---|
| ColD-Fusion2022.12 | 91.31 | |
| Multitask2022.12 | 91.27 | |
| Finetune2022.12 | 91 | |
| DSSBackbone=T5-base2024.03 | 89.51 | |
| MI-based distillationBackbone=T5-base2024.03 | 89.5 | |
| Single-taskBackbone=T5-base2024.03 | 88.88 | |
| Self-consistencyPrompting strategy=Self-consistency, Backbone model=PaLM-540B2022.03 | 88.4 | |
| FinetuningBackbone=T5-base2024.03 | 88.38 | |
| Standard-prompting (no-rationale)Prompting strategy=no-rationale, Backbone model=PaLM-540B2022.03 | 85.8 | |
| CoT-promptingPrompting strategy=Chain-of-Thought, Backbone model=PaLM-540B2022.03 | 81 | |
| text-davinci-002Prompting Strategy=E-P2022.05 | 75.6 | |
| text-davinci-002Prompting Strategy=P-E2022.05 | 69.4 | |
| text-davinci-002Prompting Strategy=FEW-SHOT2022.05 | 69.1 | |
| P-E+EXPLCALNumber of labels (L)=128, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 68.5 | |
| P-E+EXPLCALNumber of labels (L)=96, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 67.6 | |
| P-E+ZHANGNumber of labels (L)=128, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 65.9 | |
| P-E+EXPLCALNumber of labels (L)=64, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 65.8 | |
| P-E+PROBCALNumber of labels (L)=64, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 65.4 | |
| P-E+PROBCALNumber of labels (L)=96, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 65.4 | |
| P-E+PROBCALNumber of labels (L)=128, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 65.4 | |
| P-E+ZHANGNumber of labels (L)=96, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 65.4 | |
| P-E+ZHANGNumber of labels (L)=64, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 65.2 | |
| P-E+PROBCALNumber of labels (L)=32, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 64.4 | |
| P-E+EXPLCALNumber of labels (L)=32, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 64.2 | |
| FEW-SHOT+PROBCALNumber of labels (L)=128, Explanation usage=w/o Explanation2022.05 | 63.9 | |
| FEW-SHOT+PROBCALNumber of labels (L)=96, Explanation usage=w/o Explanation2022.05 | 63.2 | |
| P-E+ZHANGNumber of labels (L)=32, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 63 | |
| FEW-SHOT+PROBCALNumber of labels (L)=64, Explanation usage=w/o Explanation2022.05 | 62.4 | |
| FEW-SHOT+PROBCALNumber of labels (L)=32, Explanation usage=w/o Explanation2022.05 | 61.9 | |
| P-ENumber of labels (L)=32, Number of explanations (E)=32, Explanation usage=w/ Explanation2022.05 | 59.4 | |
| InstructGPTPrompting Strategy=P-E2022.05 | 59.4 | |
| FEW-SHOT(NN)Number of labels (L)=128, Explanation usage=w/o Explanation2022.05 | 58.9 | |
| FEW-SHOTNumber of labels (L)=32, Explanation usage=w/o Explanation2022.05 | 56.8 | |
| InstructGPTPrompting Strategy=FEW-SHOT2022.05 | 56.8 | |
| ROBERTaNumber of labels (L)=128, Explanation usage=w/o Explanation2022.05 | 54.9 | |
| MUPPET2022.12 | 52.59 | |
| ROBERTaNumber of labels (L)=96, Explanation usage=w/o Explanation2022.05 | 49 | |
| GPT-3Prompting Strategy=P-E2022.05 | 48.7 | |
| OPT (175B)Prompting Strategy=FEW-SHOT2022.05 | 44 | |
| OPT (175B)Prompting Strategy=P-E2022.05 | 43.4 | |
| GPT-3Prompting Strategy=FEW-SHOT2022.05 | 43.3 | |
| ROBERTaNumber of labels (L)=64, Explanation usage=w/o Explanation2022.05 | 43 | |
| InstructGPTPrompting Strategy=E-P2022.05 | 41.8 | |
| GPT-3Prompting Strategy=E-P2022.05 | 40.4 | |
| ROBERTaNumber of labels (L)=32, Explanation usage=w/o Explanation2022.05 | 40.1 | |
| OPT (175B)Prompting Strategy=E-P2022.05 | 39.3 |