Commonsense Question Answering on WinoGrande (WG) (val)
78.3AccuracyDeBERTa-v3-L (CANDLE Distilled)
Evaluation Results
| Method | Links | |
|---|---|---|
| DeBERTa-v3-L (CANDLE Distilled)Zero-shot=true, CSKB=CANDLE2024.01 | 78.3 | |
| CAR-DeBERTa-v3-LZero-shot=true, CSKB=AbsATM2024.01 | 78.2 | |
| CAR-DeBERTa-v3-LZero-shot=true, CSKB=ATOMIC2024.01 | 78.1 | |
| GPT-4Zero-shot=true, Variant=gpt-42024.01 | 77 | |
| DeBERTa-v3-L (MR)Zero-shot=true, CSKB=ATOMIC2024.01 | 76 | |
| Mistral-v0.1Zero-shot=true, Model Scale=7B2024.01 | 75.3 | |
| LLAMA2Zero-shot=true, Model Scale=13B2024.01 | 72.8 | |
| DeBERTa-v3-L (MR)Zero-shot=true, CSKB=ATM10X2024.01 | 71.7 | |
| VERA-T5-xxl (CANDLE Distilled)Zero-shot=true, CSKB=CANDLE2024.01 | 71.3 | |
| LLAMA2Zero-shot=true, Model Scale=7B2024.01 | 69.2 | |
| VERA-T5-xxlZero-shot=true, CSKB=AbsATM2024.01 | 68.1 | |
| VERA-T5-xxlZero-shot=true, CSKB=ATOMIC2024.01 | 67.5 | |
| VERA-T5-xxlZero-shot=true, CSKB=ATM10X2024.01 | 67.2 | |
| ChatGPT + Self-consistent chain-of-thoughtZero-shot=true, Variant=gpt-3.5-turbo, Prompting Strategy=Self-consistent chain-of-thought2024.01 | 64.1 | |
| ChatGPT + Chain-of-thoughtZero-shot=true, Variant=gpt-3.5-turbo, Prompting Strategy=Chain-of-thought2024.01 | 63.6 | |
| ChatGPTZero-shot=true, Variant=gpt-3.5-turbo2024.01 | 62.8 | |
| GPT 3.5Zero-shot=true, Variant=text-davinci-0032024.01 | 60.7 | |
| STL-AdapterZero-shot=true, CSKB=ATOMIC2024.01 | 60.3 | |
| ROBERTa-LZero-shot=true2024.01 | 57.5 | |
| Self-talkZero-shot=true2024.01 | 54.7 | |
| DeBERTa-v3-LZero-shot=true2024.01 | 50.3 |