Knowledge Base Question Answering on WebQuestion
86.02Hits@1KG-Reasoner
Evaluation Results
| Method | Links | |
|---|---|---|
| KG-ReasonerSize=30B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 86.02 | |
| PoG with GPT-4Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 84.7 | |
| GraSP2026.04 | 84.3 | |
| KBQA-o12026.04 | 82.5 | |
| KBQA-o1Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 81.6 | |
| iQUEST2026.04 | 81.2 | |
| GPT-4o-mini + KGMethodology=KG + LLMs w/o Fine-Tuning2026.04 | 81.02 | |
| LMP2026.04 | 80.4 | |
| DeepSeek-R1-Distill-Llama-70B + KGSize=70B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 75.8 | |
| ToG2026.04 | 72.8 | |
| LLaMA-3.3-70B + KGSize=70B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 72.6 | |
| FlexKBQAMethodology=KG + LLM with Fine-Tuning (Small Models Fine-Tuning)2026.04 | 68.9 | |
| DeepSeek-R1-Distill-Llama-70BSize=70B, Methodology=LLM w/o KG2026.04 | 68.84 | |
| KG-CoT with GPT 4Methodology=KG + LLM with Fine-Tuning (Small Models Fine-Tuning)2026.04 | 68 | |
| GPT-4oMethodology=LLM w/o KG2026.04 | 64.79 | |
| Qwen-2.5-7B (SFT) + KGSize=7B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 61.42 | |
| LLaMA-3.1-8B (SFT) + KGSize=8B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 60 | |
| LLaMA-3.3-70BSize=70B, Methodology=LLM w/o KG2026.04 | 59.73 | |
| ToG with GPT 4Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 57.9 | |
| GPT-4o-miniMethodology=LLM w/o KG2026.04 | 57.26 | |
| LLaMA-3.1-8B + KGSize=8B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 57.05 | |
| Qwen-2.5-7B + KGSize=7B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 56.33 | |
| LLaMA-3.1-8BSize=8B, Methodology=LLM w/o KG2026.04 | 45.88 | |
| Qwen-2.5-7BSize=7B, Methodology=LLM w/o KG2026.04 | 44.23 |