Knowledge Base Question Answering on WebQSP
96Hits@1Logits-to-Logic
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Logits-to-LogicModel Scale Category=13b model, Backbone=LLaMA3-13b2025.11 | 96 | — | — | 75.2 | — | |
| Logits-to-LogicModel Scale Category=7-8b model, Backbone=LLaMA3-8b2025.11 | 95.4 | — | — | 63.3 | — | |
| KG-ReasonerSize=30B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 93.15 | — | — | — | — | |
| GCRModel Category=Hybrid System, Backbone=LLaMA3-8b+ChatGPT2025.11 | 92.6 | — | — | 73.2 | — | |
| GCRModel Category=Hybrid System, Backbone=LLaMA3-8b+GPT4o-mini2025.11 | 92.2 | — | — | 74.1 | — | |
| RoGModel Scale Category=13b model, Backbone=LLaMA3-13b2025.11 | 89.1 | — | — | 73.6 | — | |
| STARSetting=Lightweight Retriever, RT=0.92026.04 | 88.7 | — | — | 74.1 | — | |
| RD-PSetting=Lightweight Retriever, RT=0.82026.04 | 88.1 | — | — | 73.1 | — | |
| PoG with GPT-4Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 87.3 | — | — | — | — | |
| RoGSetting=LLM-centric Retriever, RT=1.62026.04 | 87.1 | — | — | 71.6 | — | |
| RoGSize=7B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 85.7 | — | — | — | — | |
| RoGModel Scale Category=7-8b model, Backbone=LLaMA2-Chat-7B2025.11 | 85.7 | — | — | 70.8 | — | |
| Graph CoTSetting=LLM-centric Retriever, RT=123.32026.04 | 85.3 | — | — | 73.5 | — | |
| KG-CoT with GPT 4Methodology=KG + LLM with Fine-Tuning (Small Models Fine-Tuning)2026.04 | 84.9 | — | — | — | — | |
| GPT-4o-mini + KGMethodology=KG + LLMs w/o Fine-Tuning2026.04 | 84.4 | — | — | — | — | |
| LightPROF (LLaMA3-8B)Size=8B, Methodology=KG + LLM with Fine-Tuning (Small Models Fine-Tuning)2026.04 | 83.8 | — | — | — | — | |
| KG-AgentMethodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 83.3 | — | — | — | — | |
| ToG with GPT 4Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 82.6 | — | — | — | — | |
| LLaMA-3.3-70B + KGSize=70B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 81.9 | — | — | — | — | |
| DeepSeek-R1-Distill-Llama-70B + KGSize=70B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 81.34 | — | — | — | — | |
| SubgraphRAGModel Scale Category=7-8b model, Backbone=Llama3.1-8B2025.11 | 81.2 | — | — | 67.9 | — | |
| GNN-RAGModel Category=Hybrid System, Backbone=LLaMA2-7b+GNN2025.11 | 80.6 | — | — | 71.3 | — | |
| Readi with GPT-4Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 78.7 | — | — | — | — | |
| Chain-of-QuestionMethodology=KG + LLM with Fine-Tuning (Small Models Fine-Tuning)2026.04 | 78.1 | — | — | — | — | |
| DeepSeek-R1-Distill-Llama-70BSize=70B, Methodology=LLM w/o KG2026.04 | 77.43 | — | — | — | — | |
| DeepSeek-R1-70BSetting=LLM without KG, RT=16.32026.04 | 75.2 | — | — | 59 | — | |
| Qwen-2.5-7B (SFT) + KGSize=7B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 74.8 | — | — | — | — | |
| G-RetrieverSize=7B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 73.79 | — | — | — | — | |
| ChatGPT+CoTModel Category=Closed-source business model2025.11 | 73.5 | — | — | 38.5 | — | |
| LLaMA-3.1-8B (SFT) + KGSize=8B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 72.98 | — | — | — | — | |
| Qwen-2.5-7B + KGSize=7B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 72.6 | — | — | — | — | |
| GPT-4oMethodology=LLM w/o KG2026.04 | 72.55 | — | — | — | — | |
| LLaMA-3.1-8B + KGSize=8B, Methodology=KG + LLMs w/o Fine-Tuning2026.04 | 71.32 | — | — | — | — | |
| LLaMA-3.3-70BSize=70B, Methodology=LLM w/o KG2026.04 | 71.12 | — | — | — | — | |
| KD-CoTModel Scale Category=7-8b model, Backbone=LLaMA2-7b2025.11 | 68.6 | — | — | 52.5 | — | |
| ChatGPT+Few-shotModel Category=Closed-source business model2025.11 | 68.5 | — | — | 38.1 | — | |
| ToGSetting=LLM-centric Retriever, RT=41.72026.04 | 67.6 | — | — | 61.3 | — | |
| GPT-4o-miniMethodology=LLM w/o KG2026.04 | 65.97 | — | — | — | — | |
| G-RetrieverSetting=Lightweight Retriever, RT=8.92026.04 | 63.1 | — | — | 69.3 | — | |
| Llama3.3-70BSetting=LLM without KG, RT=2.82026.04 | 62.5 | — | — | 43.2 | — | |
| GRAGSetting=Lightweight Retriever, RT=10.12026.04 | 60.5 | — | — | 52.8 | — | |
| ChatGPTModel Category=Closed-source business model2025.11 | 59.3 | — | — | 43.5 | — | |
| DALKSetting=Lightweight Retriever, RT=24.72026.04 | 58.9 | — | — | 50.2 | — | |
| Qwen2.5-72BSetting=LLM without KG, RT=3.52026.04 | 58.1 | — | — | 39.4 | — | |
| LLaMA2-7bModel Scale Category=7-8b model2025.11 | 56.4 | — | — | 36.5 | — | |
| Interactive-KBQAModel Scale Category=13b model, Backbone=LLaMA2-13b2025.11 | 56.2 | — | — | 54.8 | — | |
| LLaMA3.1-8bModel Scale Category=7-8b model2025.11 | 55.5 | — | — | 34.8 | — | |
| Interactive-KBQA (13B)Size=13B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 54.86 | — | — | — | — | |
| Qwen2-7bModel Scale Category=7-8b model2025.11 | 50.8 | — | — | 35.5 | — | |
| Qwen-2.5-7BSize=7B, Methodology=LLM w/o KG2026.04 | 46.97 | — | — | — | — | |
| LLaMA-3.1-8BSize=8B, Methodology=LLM w/o KG2026.04 | 45.07 | — | — | — | — | |
| Interactive-KBQAModel Scale Category=7-8b model, Backbone=Mistral-7B2025.11 | 45 | — | — | 43.5 | — | |
| Interactive-KBQA (7B)Size=7B, Methodology=KG + LLM with Fine-Tuning (LLM Backbone Fine-Tuning)2026.04 | 43.57 | — | — | — | — | |
| CBR-KBQASupervision=Supervised2021.04 | — | 73.1 | 75.1 | 72.8 | 69.9 | |
| CoT2026.01 | — | — | — | — | 62.2 | |
| DoG2026.01 | — | — | — | — | 91 | |
| DP2026.01 | — | — | — | — | 87.5 | |
| EmbedKGQASupervision=Weakly supervised2021.04 | — | — | — | 66.6 | — | |
| FLAREFramework=ToG2026.01 | — | — | — | — | 90.4 | |
| FLAREFramework=PoG2026.01 | — | — | — | — | 93.9 | |
| GCR2026.01 | — | — | — | — | 92.2 | |
| GraftNetSupervision=Weakly supervised2021.04 | — | — | — | 66.4 | — | |
| GSRModel Category=Hybrid System, Backbone=T5-3b + LLaMA2-7b2025.11 | — | — | — | 58.9 | — | |
| IO Prompt2026.01 | — | — | — | — | 63.3 | |
| KB-BINDER2026.01 | — | — | — | — | 74.4 | |
| LATS2026.01 | — | — | — | — | 76.5 | |
| PathMind2026.01 | — | — | — | — | 89.5 | |
| PoG2026.01 | — | — | — | — | 87.3 | |
| ProgRAG2026.01 | — | — | — | — | 90.4 | |
| PullNetSupervision=Weakly supervised2021.04 | — | — | — | 68.1 | — | |
| RAP2026.01 | — | — | — | — | 75.6 | |
| RoG2026.01 | — | — | — | — | 85.7 | |
| RwT2026.01 | — | — | — | — | 87 | |
| SC2026.01 | — | — | — | — | 61.1 | |
| STAGGSupervision=Supervised2021.04 | — | 70.9 | 80.3 | 71.7 | 63.9 | |
| T5-11BSupervision=Supervised2021.04 | — | 62.1 | 62.6 | 61.5 | 56.5 | |
| T5-11B + ReviseSupervision=Supervised2021.04 | — | 63.6 | 64.3 | 63 | 57.7 | |
| T5-3BBackbone=T5-3B2022.01 | — | — | — | 80.7 | — | |
| T5-baseBackbone=T5-base2022.01 | — | — | — | 78.83 | — | |
| T5-largeBackbone=T5-large2022.01 | — | — | — | 79.45 | — | |
| ToG2026.01 | — | — | — | — | 82.6 | |
| ToG 2.02026.01 | — | — | — | — | 81.1 | |
| ToT2026.01 | — | — | — | — | 73.5 | |
| Ye et al. (2021b)Extra Pre-training=false2022.01 | — | — | — | 83.6 | — |