Multi-Hop Knowledge Graph Question Answering on WebQSP
96.7Hits@1PoG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PoGClass=ICL, LLM=GPT-4, External Knowledge=With external knowledge2024.10 | 96.7 | — | |
| PoG-EClass=ICL, LLM=GPT-4, External Knowledge=With external knowledge2024.10 | 95.4 | — | |
| Logits-to-LogicBackbone LLM=LLaMA3.1-8b2025.11 | 95.4 | — | |
| PoGClass=ICL, LLM=GPT-3.5-Turbo, External Knowledge=With external knowledge2024.10 | 93.9 | — | |
| GCRMethod Paradigm=Agentic Reasoning, Backbone LLM=LLaMA3.1-8b2025.11 | 92.2 | — | |
| DoGMethod Paradigm=Agentic Reasoning2025.11 | 91 | — | |
| PoG-EClass=ICL, LLM=GPT-3.5-Turbo, External Knowledge=With external knowledge2024.10 | 90.9 | — | |
| RSF-GLLMCategory=Ours, Backbone=Qwen3-8B2026.07 | 90.45 | 79.15 | |
| CoGBackbone=GPT-4, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 89.7 | — | |
| RSF-GLLMCategory=Ours, Backbone=LLaMA2-7B2026.07 | 89.5 | 78.93 | |
| FD-PORTCategory=Agentic Search, Backbone=LLaMA3-8B2026.07 | 89.2 | — | |
| iQUESTBackbone=GPT-4o2025.06 | 88.93 | — | |
| PoGBackbone=GPT-4, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 87.3 | — | |
| CoGBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 86.8 | — | |
| PrivGemoType=Proposed, Brain LLM=GPT-3.5-Turbo, Hand LLM=Qwen3-32b2026.01 | 86 | — | |
| PrivGemoType=Proposed, Brain LLM=DeepSeek-V3, Hand LLM=Qwen3-32b2026.01 | 85.9 | — | |
| Prior FT SOTAClass=SL, External Knowledge=With external knowledge2024.10 | 85.7 | — | |
| RoGEvaluation Protocol=Fine-Tuning, KG-Augmented=true2026.01 | 85.7 | — | |
| RoGMethod Paradigm=Agentic Reasoning2025.11 | 85.7 | — | |
| RoGCategory=LLM + KG, Backbone=LLaMA2-7B2026.07 | 85.7 | 70.8 | |
| KG-CoT2025.06 | 84.9 | — | |
| KG-CoTMethod Paradigm=Agentic Reasoning, Backbone LLM=ChatGPT2025.11 | 84.9 | — | |
| GoGMethod Paradigm=Agentic Reasoning2025.11 | 84.4 | — | |
| PrivGemoType=Proposed, Brain LLM=GPT-4o-mini, Hand LLM=Qwen3-32b2026.01 | 84.3 | — | |
| KG-AgentEvaluation Protocol=Fine-Tuning, KG-Augmented=true2026.01 | 83.3 | — | |
| KG-AgentMethod Paradigm=Agentic Reasoning2025.11 | 83.3 | — | |
| EffiQACategory=Agentic Search, Backbone=Llama3.1-8B2026.07 | 82.9 | — | |
| GNN-RAG + RACategory=LLM + KG, Backbone=LLaMA2-7B2026.07 | 82.8 | 73.5 | |
| ToG/ToG-RClass=ICL, LLM=GPT-4, External Knowledge=With external knowledge2024.10 | 82.6 | — | |
| ToGBackbone=GPT-4, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 82.6 | — | |
| ToGType=KG-centric RAG, LLM=GPT-42026.01 | 82.6 | — | |
| ToG2025.06 | 82.6 | — | |
| ToGMethod Paradigm=Agentic Reasoning, Backbone LLM=GPT42025.11 | 82.6 | — | |
| DECAFEvaluation Protocol=Fine-Tuning, KG-Augmented=true2026.01 | 82.1 | — | |
| DECAFCategory=LLM + KG, Backbone=FiD-3B2026.07 | 82.1 | 78.8 | |
| PoGBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 82 | — | |
| PoGType=KG-centric RAG, LLM=GPT-3.5-Turbo2026.01 | 82 | — | |
| PoGMethod Paradigm=Agentic Reasoning, Backbone LLM=ChatGPT2025.11 | 82 | — | |
| ToG-RMethod Paradigm=Agentic Reasoning, Backbone LLM=GPT42025.11 | 81.9 | — | |
| ToG-2.0Class=ICL, LLM=GPT-3.5-Turbo, External Knowledge=With external knowledge2024.10 | 81.1 | — | |
| ToG-2Type=KG-centric RAG, LLM=GPT-3.5-Turbo2026.01 | 81.1 | — | |
| GNN-RAGCategory=LLM + KG, Backbone=LLaMA2-7B2026.07 | 80.6 | 71.3 | |
| UniKGQAEvaluation Protocol=Fine-Tuning, KG-Augmented=true2026.01 | 79.1 | — | |
| GoGType=KG-centric RAG, LLM=GPT-3.5-Turbo2026.01 | 78.7 | — | |
| SymAgentMethod Paradigm=Agentic Reasoning2025.11 | 78.5 | — | |
| Chain-of-QuestionLLM=GPT-3.5-turbo, Question Decomposition Model=T52025.06 | 78.1 | — | |
| UniKGQACategory=Graph Retrieval2026.07 | 77.2 | 72.2 | |
| ReaRevCategory=Graph Retrieval2026.07 | 76.4 | 70.9 | |
| ToG/ToG-RClass=ICL, LLM=GPT-3.5-Turbo, External Knowledge=With external knowledge2024.10 | 76.2 | — | |
| ToGBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 76.2 | — | |
| ToGType=KG-centric RAG, LLM=GPT-3.5-Turbo2026.01 | 76.2 | — | |
| ToGMethod Paradigm=Agentic Reasoning, Backbone LLM=ChatGPT2025.11 | 76.2 | — | |
| ToG-RMethod Paradigm=Agentic Reasoning, Backbone LLM=ChatGPT2025.11 | 75.8 | — | |
| ARoGType=Privacy-aware KGQA, LLM=GPT-4o-mini2026.01 | 74.7 | — | |
| RE-KBQAEvaluation Protocol=Fine-Tuning, KG-Augmented=true2026.01 | 74.6 | — | |
| CoGBackbone=Qwen2.5-7B, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 74.5 | — | |
| KB-BINDERClass=ICL, LLM=Codex, External Knowledge=With external knowledge2024.10 | 74.4 | — | |
| KB-BINDERBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 74.4 | — | |
| KD-CoTBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 73.7 | — | |
| KD-CoTMethod Paradigm=Agentic Reasoning2025.11 | 73.7 | — | |
| StructGPTBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 72.6 | — | |
| StructGPTMethod Paradigm=Agentic Reasoning2025.11 | 72.6 | — | |
| TransferNetCategory=Embedding-based2026.07 | 71.4 | — | |
| Interactive-KBQA2025.06 | 71.2 | — | |
| G-RetreiverCategory=LLM + KG2026.07 | 70.1 | — | |
| SR+NSMCategory=Graph Retrieval2026.07 | 68.9 | 64.1 | |
| NSMCategory=Embedding-based2026.07 | 68.7 | 62.8 | |
| KD-CoTCategory=LLM + KG2026.07 | 68.6 | 52.5 | |
| PullNetCategory=Graph Retrieval2026.07 | 68.1 | — | |
| EmbedKGQACategory=Embedding-based2026.07 | 66.6 | — | |
| GraftNetCategory=Graph Retrieval2026.07 | 66.4 | 60.4 | |
| IO promptLLM=GPT-3.5-Turbo, External Knowledge=Without external knowledge2024.10 | 63.3 | — | |
| IO PromptBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=false2026.01 | 63.3 | — | |
| IO promptType=LLM-only, LLM=GPT-3.5-Turbo2026.01 | 63.3 | — | |
| IO promptMethod Paradigm=LLMs Reasoning, Backbone LLM=ChatGPT2025.11 | 63.3 | — | |
| CoTLLM=GPT-3.5-Turbo, External Knowledge=Without external knowledge2024.10 | 62.2 | — | |
| CoTBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=false2026.01 | 62.2 | — | |
| CoTType=LLM-only, LLM=GPT-3.5-Turbo2026.01 | 62.2 | — | |
| CoTMethod Paradigm=LLMs Reasoning, Backbone LLM=ChatGPT2025.11 | 62.2 | — | |
| SCLLM=GPT-3.5-Turbo, External Knowledge=Without external knowledge2024.10 | 61.1 | — | |
| SCBackbone=GPT-3.5, Evaluation Protocol=Prompting, KG-Augmented=false2026.01 | 61.1 | — | |
| SCType=LLM-only, LLM=GPT-3.5-Turbo2026.01 | 61.1 | — | |
| SCMethod Paradigm=LLMs Reasoning, Backbone LLM=ChatGPT2025.11 | 61.1 | — | |
| PoGBackbone=Qwen2.5-7B, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 58.5 | — | |
| ToGBackbone=Qwen2.5-7B, Evaluation Protocol=Prompting, KG-Augmented=true2026.01 | 56 | — | |
| LLaMA3.1-8BCategory=LLM Reasoning, Zero-shot=true2026.07 | 55.1 | 35.6 | |
| Qwen3-8BCategory=LLM Reasoning, Zero-shot=true2026.07 | 50.1 | 34 | |
| KV-MemCategory=Embedding-based2026.07 | 46.7 | 34.5 | |
| FlexKBQA2025.06 | 46.2 | — | |
| DARAMethod Paradigm=Agentic Reasoning, Backbone LLM=LLaMA2-13b2025.11 | 30.3 | — |