Question Answering on CommonsenseQA IH (test)
88.9AccuracyHuman Performance
Evaluation Results
| Method | Links | |
|---|---|---|
| Human PerformanceTraining Data Ratio=10%2019.09 | 88.9 | |
| Human PerformanceTraining Data Ratio=50%2019.09 | 88.9 | |
| Human PerformanceTraining Data Ratio=100%2019.09 | 88.9 | |
| KnowGPTCategory=Ours2023.12 | 81.8 | |
| GPT-4Category=LLM + Zero-shot2023.12 | 78.6 | |
| MindmapCategory=LLM + KG Prompting2023.12 | 78.4 | |
| Ours2023.05 | 74.97 | |
| GrapeQACategory=KG-enhanced LM2023.12 | 74.9 | |
| GSC2023.05 | 74.48 | |
| JointLKCategory=KG-enhanced LM2023.12 | 74.4 | |
| GreaseLM2023.05 | 74.2 | |
| GreaseLMCategory=KG-enhanced LM2023.12 | 74.2 | |
| SAFE2023.05 | 74.03 | |
| HamQACategory=KG-enhanced LM2023.12 | 73.9 | |
| CoKCategory=LLM + KG Prompting2023.12 | 73.9 | |
| QA-GNN2023.05 | 73.41 | |
| RoGCategory=LLM + KG Prompting2023.12 | 73.4 | |
| QA-GNNCategory=KG-enhanced LM2023.12 | 73.3 | |
| Llama3Category=LLM + Zero-shot, Parameters=8b2023.12 | 72.3 | |
| MHGRNCategory=KG-enhanced LM2023.12 | 71.3 | |
| MHGRN2023.05 | 71.11 | |
| GPT-3.5Category=LLM + Zero-shot, Implementation=gpt-3.5-turbo2023.12 | 71 | |
| RN2023.05 | 69.08 | |
| KagNet2023.05 | 69.01 | |
| RoBerta-largeCategory=LM + Fine tuning2023.12 | 68.7 | |
| RoBERTa-largeKnowledge Graph=without2023.05 | 68.69 | |
| GconAttn2023.05 | 68.59 | |
| RGCN2023.05 | 68.41 | |
| BERT-LARGE-KAGNETTraining Data Ratio=100%, Backbone=BERT-Large2019.09 | 57.16 | |
| BERT-BASE-KAGNETTraining Data Ratio=100%, Backbone=BERT-Base2019.09 | 56.19 | |
| BERT-LARGE-FINETUNINGTraining Data Ratio=100%, Backbone=BERT-Large2019.09 | 55.84 | |
| Bert-largeCategory=LM + Fine tuning2023.12 | 55.4 | |
| Llama2Category=LLM + Zero-shot, Parameters=7b2023.12 | 54.6 | |
| Bert-baseCategory=LM + Fine tuning2023.12 | 53.5 | |
| BERT-BASE-FINETUNINGTraining Data Ratio=100%, Backbone=BERT-Base2019.09 | 53.26 | |
| GPT-3Category=LLM + Zero-shot, Implementation=text-davinci-0022023.12 | 52 | |
| BERT-LARGE-KAGNETTraining Data Ratio=50%, Backbone=BERT-Large2019.09 | 51.13 | |
| BERT-LARGE-FINETUNINGTraining Data Ratio=50%, Backbone=BERT-Large2019.09 | 49.88 | |
| BaichuanCategory=LLM + Zero-shot, Parameters=7B2023.12 | 47.6 | |
| ChatGLMCategory=LLM + Zero-shot2023.12 | 46.9 | |
| GPT-KAGNETTraining Data Ratio=100%, Backbone=GPT2019.09 | 46.79 | |
| GPT-FINETUNINGTraining Data Ratio=100%, Backbone=GPT2019.09 | 45.58 | |
| InternLMCategory=LLM + Zero-shot2023.12 | 45.4 | |
| ChatGLM2Category=LLM + Zero-shot2023.12 | 42.5 | |
| BERT-BASE-KAGNETTraining Data Ratio=50%, Backbone=BERT-Base2019.09 | 39.01 | |
| BERT-BASE-FINETUNINGTraining Data Ratio=50%, Backbone=BERT-Base2019.09 | 36.83 | |
| BERT-LARGE-KAGNETTraining Data Ratio=10%, Backbone=BERT-Large2019.09 | 33.91 | |
| BERT-LARGE-FINETUNINGTraining Data Ratio=10%, Backbone=BERT-Large2019.09 | 32.88 | |
| GPT-KAGNETTraining Data Ratio=50%, Backbone=GPT2019.09 | 32.33 | |
| GPT-FINETUNINGTraining Data Ratio=50%, Backbone=GPT2019.09 | 31.28 | |
| BERT-BASE-KAGNETTraining Data Ratio=10%, Backbone=BERT-Base2019.09 | 30.94 | |
| BERT-BASE-FINETUNINGTraining Data Ratio=10%, Backbone=BERT-Base2019.09 | 29.78 | |
| GPT-KAGNETTraining Data Ratio=10%, Backbone=GPT2019.09 | 26.98 | |
| GPT-FINETUNINGTraining Data Ratio=10%, Backbone=GPT2019.09 | 26.51 | |
| Random guessTraining Data Ratio=10%2019.09 | 20 | |
| Random guessTraining Data Ratio=50%2019.09 | 20 | |
| Random guessTraining Data Ratio=100%2019.09 | 20 |