Commonsense Reasoning on CommonsenseQA (dev)
73.58Accuracy (5%)T5-3B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| T5-3BAugmentation=VAWI-SBS+PET2022.12 | 73.58 | 78.58 | 79.66 | 84.56 | — | |
| T5-3BAugmentation=+Images2022.12 | 72.87 | 76.17 | 78.71 | 83.64 | — | |
| T5-3BAugmentation=+None2022.12 | 71.99 | 75.27 | 77.72 | 82.4 | — | |
| RoBERTa-largeAugmentation=+VAWI-SBS2022.12 | 52.98 | 61.97 | 67.4 | 79.19 | — | |
| RoBERTa-largeAugmentation=+Images2022.12 | 52.18 | 60.93 | 66.08 | 78.39 | — | |
| RoBERTa-largeAugmentation=+None2022.12 | 51.24 | 59.95 | 65.52 | 76.65 | — | |
| RoBERTa-baseAugmentation=+VAWI-SBS2022.12 | 46.51 | 52.44 | 59.87 | 71.01 | — | |
| RoBERTa-baseAugmentation=+Images2022.12 | 45.72 | 51.17 | 58.96 | 69.64 | — | |
| RoBERTa-baseAugmentation=+None2022.12 | 44.88 | 50.04 | 57.08 | 67.9 | — | |
| DeBERTaV3-large + KEARScenario=Scenario 32023.11 | — | — | — | — | 91.2 | |
| GPT-3.5-turboPrompting Strategy=Few-Shot CoT, Scenario=Scenario 32023.11 | — | — | — | — | 74.5 | |
| GPT-3.5-turboPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | — | — | — | — | 79 | |
| GPT-3.5-turboPrompting Strategy=USC, Scenario=Scenario 32023.11 | — | — | — | — | 48.9 | |
| LaMDAPrompting Strategy=Few-Shot CoT, Scenario=Scenario 32023.11 | — | — | — | — | 57.9 | |
| LaMDAPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | — | — | — | — | 63.1 | |
| PaLMPrompting Strategy=Few-Shot CoT, Scenario=Scenario 32023.11 | — | — | — | — | 79 | |
| PaLMPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | — | — | — | — | 80.7 | |
| Self-AgreementModel=GPT-3.5-turbo, Scenario=Scenario 32023.11 | — | — | — | — | 79.4 |