Knowledge Graph Question Answering on GrailQA (test)
91Overall ScoreCoG
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CoGMethod Category=Prompting, Backbone LLM=DeepSeek-V3.22026.06 | 91 | 90.5 | 84.2 | 93.5 | |
| CoGMethod Category=Prompting, Backbone LLM=GPT-4.1-mini2026.06 | 89.4 | 89 | 83.6 | 91.6 | |
| PoGSetup=Prompting KG-Augmented LLM, Underlying LLM=GPT-42024.10 | 84.7 | 87.9 | 69.7 | 88.6 | |
| CoGMethod Category=Prompting, Backbone LLM=Qwen3-Coder-30B-A3B2026.06 | 82 | 82.5 | 77.9 | 83.2 | |
| ToGSetup=Prompting KG-Augmented LLM, Underlying LLM=GPT-42024.10 | 81.4 | 79.4 | 67.3 | 86.5 | |
| ReKnoSMethod Category=Prompting, Backbone LLM=GPT-4o-mini2026.06 | 80.5 | — | — | — | |
| SRPMethod Category=Prompting, Backbone LLM=GPT-4.1-mini2026.06 | 78.8 | 75.8 | 62.6 | 85.8 | |
| PoGMethod Category=Prompting, Backbone LLM=Qwen3-Coder-30B-A3B2026.06 | 76.8 | 77.9 | 61.1 | 81.9 | |
| PoGSetup=Prompting KG-Augmented LLM, Underlying LLM=GPT-3.5 or others2024.10 | 76.5 | 76.3 | 62.1 | 81.7 | |
| GAINSetup=Fine-Tuned KG-Augmented LLM2024.10 | 76.3 | 88.5 | 73.7 | 71.8 | |
| PanguSetup=Fine-Tuned KG-Augmented LLM2024.10 | 75.4 | 84.4 | 74.6 | 71.6 | |
| PanguMethod Category=Fine-tuning2026.06 | 75.4 | 84.4 | 74.6 | 71.6 | |
| FC-KBQASetup=Fine-Tuned KG-Augmented LLM2024.10 | 73.2 | 88.5 | 70 | 67.6 | |
| TIARASetup=Fine-Tuned KG-Augmented LLM2024.10 | 73 | 87.8 | 69.2 | 68 | |
| TIARAMethod Category=Fine-tuning2026.06 | 73 | 87.8 | 69.2 | 68 | |
| ReadiMethod Category=Prompting, Backbone LLM=GPT-4.1-mini2026.06 | 71.7 | 67.5 | 60.1 | 77.6 | |
| PoGMethod Category=Prompting, Backbone LLM=DeepSeek-V3.22026.06 | 71.6 | 68.4 | 56.4 | 78.3 | |
| RnG-KBQASetup=Fine-Tuned KG-Augmented LLM2024.10 | 68.8 | 86.2 | 63.8 | 63 | |
| ToGSetup=Prompting KG-Augmented LLM, Underlying LLM=GPT-3.5 or others2024.10 | 68.7 | 70.1 | 56.1 | 72.7 | |
| FlexKBQASetup=Fine-Tuned KG-Augmented LLM2024.10 | 62.8 | 71.3 | 59.1 | 60.6 | |
| FlexKBQAMethod Category=Fine-tuning2026.06 | 62.8 | 71.3 | 59.1 | 60.6 | |
| KB-BINDERSetup=Prompting KG-Augmented LLM, Underlying LLM=GPT-3.5 or others2024.10 | 50.6 | — | — | — | |
| IO PromptingMethod Category=Prompting, Backbone LLM=DeepSeek-V3.22026.06 | 35.7 | 34.8 | 26 | 39.6 | |
| IO PromptingMethod Category=Prompting, Backbone LLM=Qwen3-Coder-30B-A3B2026.06 | 33.4 | 33.5 | 23 | 37.1 | |
| SCSetup=LLM-Only2024.10 | 29.6 | — | — | — | |
| IO PromptSetup=LLM-Only2024.10 | 29.4 | — | — | — | |
| CoTSetup=LLM-Only2024.10 | 28.1 | — | — | — |