Knowledge Graph Question Answering on CWQ
85.2Hit@1AGE AMAR
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AGE AMARLLM=Llama2-7B-LoRA, Non-parameter Retriever=true2026.06 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| AGE AMARLLM=Llama2-13B-LoRA, Non-parameter Retriever=true2026.06 | 85.1 | — | — | — | — | — | — | — | — | — | — | |
| AMARLLM=Llama2-13B-LoRA, Non-parameter Retriever=true2026.06 | 83.1 | — | — | — | — | — | — | — | — | — | — | |
| AMARLLM=Llama2-7B-LoRA, Non-parameter Retriever=true2026.06 | 82.9 | — | — | — | — | — | — | — | — | — | — | |
| DualRLLM=ChatGPT, Trainable Retriever GNN=true, Trainable Retriever LLM=true2026.06 | 82.8 | — | — | — | — | — | — | — | — | — | — | |
| PoGLLM=ChatGPT, Trainable Retriever LLM=true2026.06 | 82 | — | — | — | — | — | — | — | — | — | — | |
| RoGLLM=ChatGPT, Trainable Retriever LLM=true2026.06 | 80 | — | — | — | — | — | — | — | — | — | — | |
| S-Path-RAG +IterativeType=S-Path-RAG2026.03 | 79.3 | 76.9 | — | — | 97.3 | — | — | — | — | — | — | |
| CoGMethod Category=Prompting, Backbone LLM=DeepSeek-V3.22026.06 | 79.1 | — | — | — | — | — | — | — | — | — | — | |
| S-Path-RAGType=S-Path-RAG2026.03 | 77.9 | 75.2 | — | — | 95.7 | — | — | — | — | — | — | |
| ToGLLM=ChatGPT, Trainable Retriever LLM=true2026.06 | 76.2 | — | — | — | — | — | — | — | — | — | — | |
| GCRLLM Backbone=GPT-4o-mini2026.01 | 75.8 | 61.7 | — | — | — | — | — | — | — | — | — | |
| GoGType=LLMs + KG2025.12 | 75.2 | — | — | — | — | — | — | — | — | — | — | |
| PoGLLM=GPT-4, Trainable Retriever LLM=true2026.06 | 75 | — | — | — | — | — | — | — | — | — | — | |
| KG-R1Type=RL+KG+LLM2026.03 | 73.8 | 70.9 | — | — | 69.7 | — | — | — | — | — | — | |
| DualRLLM=GPT-4, Trainable Retriever GNN=true, Trainable Retriever LLM=true2026.06 | 73.6 | — | — | — | — | — | — | — | — | — | — | |
| ORTType=LLM+KGs(non-Fine-tuned)2025.02 | 72.9 | 62.6 | — | — | — | — | — | — | — | — | — | |
| GCRLLM Backbone=ChatGPT2026.01 | 72.7 | 60.9 | — | — | — | — | — | — | — | — | — | |
| PoGMethod Category=Prompting, Backbone LLM=DeepSeek-V3.22026.06 | 72.6 | — | — | — | — | — | — | — | — | — | — | |
| ToGBackbone Model=GPT-42026.05 | 72.5 | — | — | — | — | — | — | — | — | — | — | |
| RPO-RAGLLM Backbone=Llama3.1–8B2026.01 | 72.3 | 64.5 | — | — | — | — | — | — | — | — | — | |
| KG-AgentType=LLMs + KG2025.12 | 72.2 | 69.8 | — | — | — | — | — | — | — | — | — | |
| KG-AgentMethod Category=Fine-tuning2026.06 | 72.2 | — | — | — | — | — | — | — | — | — | — | |
| CoGMethod Category=Prompting, Backbone LLM=GPT-4.1-mini2026.06 | 72.2 | — | — | — | — | — | — | — | — | — | — | |
| PATHISELLM=GPT-4.1, Training Strategy=Ours2026.05 | 71.9 | 61.3 | — | — | — | — | — | — | — | — | — | |
| FiDeLiSType=LLMs + KG2025.12 | 71.5 | 64.3 | — | — | — | — | — | — | — | — | — | |
| PathHDType=LLMs + KG2025.12 | 71.5 | 65.8 | — | — | — | — | — | — | — | — | — | |
| FiDeLiSLLM Backbone=gpt-4-turbo, Category=Prompting - LLM + KG2024.05 | 71.47 | 64.32 | — | — | — | — | — | — | — | — | — | |
| ParallaxRAG + GPT-4o (200)Generator=GPT-4o, Number of retrieved triples=2002025.10 | 70.74 | — | — | — | — | — | 62.31 | 66.69 | — | — | — | |
| DeCAFCategory=Finetuning - LLM + KG2024.05 | 70.42 | — | — | — | — | — | — | — | — | — | — | |
| Prior FT SOTA2026.05 | 70.4 | — | — | — | — | — | — | — | — | — | — | |
| DeCAFMethod Category=Fine-tuning2026.06 | 70.4 | — | — | — | — | — | — | — | — | — | — | |
| PATHISELLM=GPT-4o, Training Strategy=Ours2026.05 | 69.7 | 61.5 | — | — | — | — | — | — | — | — | — | |
| ToG+GPT-4Type=KG+LLM2026.03 | 69.5 | 62.6 | — | — | 83.7 | — | — | — | — | — | — | |
| ToGLLM=GPT-4, Trainable Retriever LLM=true2026.06 | 69.5 | — | — | — | — | — | — | — | — | — | — | |
| ReasoningLMCategory=Training-based Method2024.03 | 69 | — | — | — | — | — | — | — | — | — | — | |
| RAPLLLM=GPT-4o, Training Strategy=Training with LLM-refined supervision2026.05 | 69 | 58.8 | — | — | — | — | — | — | — | — | — | |
| SRPMethod Category=Prompting, Backbone LLM=GPT-4.1-mini2026.06 | 69 | — | — | — | — | — | — | — | — | — | — | |
| ReGLLM=GPT-4o, Training Strategy=Training with LLM-refined supervision2026.05 | 68.9 | 62.5 | — | — | — | — | — | — | — | — | — | |
| TOGLLM Backbone=gpt-4-turbo, Category=Prompting - LLM + KG2024.05 | 68.51 | 60.2 | — | — | — | — | — | — | — | — | — | |
| ToGLLM Backbone=GPT-42026.01 | 68.5 | — | — | — | — | — | — | — | — | — | — | |
| Think-on-GraphType=LLMs + KG2025.12 | 68.5 | 60.2 | — | — | — | — | — | — | — | — | — | |
| REL-RAG + GPT-4o-miniGenerator=GPT-4o-mini2025.10 | 68.3 | — | — | — | — | — | 58.6 | — | — | — | — | |
| ReKnoSLLM=GPT-4, Trainable Retriever GNN=true, Trainable Retriever LLM=true2026.06 | 68.2 | — | — | — | — | — | — | — | — | — | — | |
| RPO-RAGLLM Backbone=Llama2–7B2026.01 | 68.1 | 59.2 | — | — | — | — | — | — | — | — | — | |
| ToG (GPT-4)Generator=GPT-42025.10 | 67.6 | — | — | — | — | — | — | — | — | — | — | |
| SubgraphRAG (100 triples)LLM=GPT-4o, Training Strategy=Training with weakly supervised paths2026.05 | 67.3 | 59.2 | — | — | — | — | — | — | — | — | — | |
| CBR-KBQACategory=Finetuning - LLM + KG2024.05 | 67.14 | — | — | — | — | — | — | — | — | — | — | |
| Readi-GPT4Category=Inference-based Method, Backbone=GPT42024.03 | 67 | — | — | — | — | — | — | — | — | — | — | |
| GNN-RAG2025.10 | 66.8 | — | — | — | — | — | 59.4 | — | — | — | — | |
| ReKnoSMethod Category=Prompting, Backbone LLM=GPT-4o-mini2026.06 | 66.8 | — | — | — | — | — | — | — | — | — | — | |
| GNN-RAGLLM=ChatGPT, Trainable Retriever GNN=true2026.06 | 66.8 | — | — | — | — | — | — | — | — | — | — | |
| SubgraphRAG + GPT-4oGenerator=GPT-4o2025.10 | 66.69 | — | — | — | — | — | 59.08 | 66.57 | — | — | — | |
| RoECategory=LLMs+KGs, Backbone=Llama3.1-8B-Instruct2025.10 | 66.49 | 53.21 | — | 1.25 | — | — | — | — | — | — | — | |
| ParallaxRAG + Qwen3-30B (200)Generator=Qwen3-30B, Number of retrieved triples=2002025.10 | 66.21 | — | — | — | — | — | 59.31 | 58.06 | — | — | — | |
| EPERMType=LLM2026.03 | 66.2 | 58.9 | — | — | 89.7 | — | — | — | — | — | — | |
| RPO-RAGLLM Backbone=Llama3.2–3B2026.01 | 66 | 57.3 | — | — | — | — | — | — | — | — | — | |
| ReknoSBase Model=Qwen 3-235B, Modules=32025.09 | 65.63 | — | — | — | — | — | — | — | — | 2,677.92 | 684.38 | |
| KG-R1Base Model=Qwen2.5-3B, Modules=12025.09 | 65.31 | — | — | — | — | — | — | — | — | 3,205.95 | 311.14 | |
| ParallaxRAG + Qwen3-30BGenerator=Qwen3-30B, Number of retrieved triples=1002025.10 | 65.3 | — | — | — | — | — | 59.25 | 57.18 | — | — | — | |
| GNN-RAGLLM=LLaMA2-Chat-7B, Fine-tuned=true, Training Strategy=Training with weakly supervised paths2026.05 | 65.3 | 58.3 | — | — | — | — | — | — | — | — | — | |
| RoG + GraphRAG-FI2025.10 | 64.82 | — | — | — | — | — | 55.12 | — | — | — | — | |
| GCRLLM=GPT-4o, Training Strategy=Training with weakly supervised paths2026.05 | 64.6 | 57.1 | — | — | — | — | — | — | — | — | — | |
| ParallaxRAG + Qwen3-30B (500)Generator=Qwen3-30B, Number of retrieved triples=5002025.10 | 64.52 | — | — | — | — | — | 59.18 | 60.07 | — | — | — | |
| PoGMethod Category=Prompting, Backbone LLM=Qwen3-Coder-30B-A3B2026.06 | 64.1 | — | — | — | — | — | — | — | — | — | — | |
| GNNRAGCategory=LLMs+GNNs, Backbone=Llama3.1-8B-Instruct2025.10 | 63.61 | 54.92 | — | 2 | — | — | — | — | — | — | — | |
| DPLLM=GPT-4o, Training Strategy=Training with weakly supervised paths2026.05 | 63.2 | 57.3 | — | — | — | — | — | — | — | — | — | |
| SGRBackbone Model=GPT-42026.05 | 63.2 | — | 59 | — | — | — | — | — | — | — | — | |
| FiDeLiSLLM Backbone=gpt-3.5-turbo, Category=Prompting - LLM + KG2024.05 | 63.12 | 61.78 | — | — | — | — | — | — | — | — | — | |
| GNN-RAG+RAType=GNN+LLM2026.03 | 62.8 | 60.4 | — | — | 79.3 | — | — | — | — | — | — | |
| RoGType=LLM+KGs(Fine-tuned)2025.02 | 62.6 | 56.2 | — | — | — | — | — | — | — | — | — | |
| ROGCategory=Training-based Method2024.03 | 62.6 | — | — | — | — | — | — | — | — | — | — | |
| ROGType=LLMs + KG2025.12 | 62.6 | 56.2 | — | — | — | — | — | — | — | — | — | |
| RoGMethod Category=Fine-tuning2026.06 | 62.6 | — | — | — | — | — | — | — | — | — | — | |
| NeuroSymActiveBackbone=LLaMa3-8B2026.02 | 62.5 | — | — | — | — | — | — | — | — | — | — | |
| SubgraphRAGLLM Backbone=GPT-4o-mini2026.01 | 62 | 54.1 | — | — | — | — | — | — | — | — | — | |
| EtD2025.10 | 62 | — | — | — | — | — | — | — | — | — | — | |
| RoG-Joint2025.10 | 61.94 | — | — | — | — | — | 54.63 | 55.15 | — | — | — | |
| GNN-RAGType=GNN+LLM2026.03 | 61.7 | 59.4 | — | — | 75.6 | — | — | — | — | — | — | |
| KG-HopperLLM=Qwen-2.5-7B, Fine-tuned=true, Training Strategy=Training without intermediate supervision2026.05 | 61.7 | — | — | — | — | — | — | — | — | — | — | |
| RoGCategory=Finetuning - LLM + KG2024.05 | 61.39 | 56.17 | — | — | — | — | — | — | — | — | — | |
| ReknoSBase Model=Qwen 2.5-72B, Modules=32025.09 | 60.88 | — | — | — | — | — | — | — | — | 2,278.71 | 548.52 | |
| CoGMethod Category=Prompting, Backbone LLM=Qwen3-Coder-30B-A3B2026.06 | 60.8 | — | — | — | — | — | — | — | — | — | — | |
| RoG-Sep2025.10 | 60.55 | — | — | — | — | — | 53.87 | 54.51 | — | — | — | |
| RPO-RAGLLM Backbone=Llama3.2–1B2026.01 | 60.3 | 50.4 | — | — | — | — | — | — | — | — | — | |
| ReadiMethod Category=Prompting, Backbone LLM=GPT-4.1-mini2026.06 | 60.2 | — | — | — | — | — | — | — | — | — | — | |
| LightPROFBackbone=LLaMa3-8B2026.02 | 59.3 | — | — | — | — | — | — | — | — | — | — | |
| PoGLLM=GPT-4o, Training Strategy=Prompting with in-context learning2026.05 | 59.3 | 51.9 | — | — | — | — | — | — | — | — | — | |
| RoGLLM=LLaMA2-Chat-7B, Fine-tuned=true, Training Strategy=Training with weakly supervised paths2026.05 | 59.2 | 53.6 | — | — | — | — | — | — | — | — | — | |
| ToG+ChatGPTType=KG+LLM2026.03 | 58.9 | 53 | — | — | 81.2 | — | — | — | — | — | — | |
| ReKnoSLLM=ChatGPT, Trainable Retriever GNN=true, Trainable Retriever LLM=true2026.06 | 58.5 | — | — | — | — | — | — | — | — | — | — | |
| ParallaxRAG + Llama3.1-8BGenerator=Llama3.1-8B, Number of retrieved triples=1002025.10 | 58.41 | — | — | — | — | — | 48.33 | 66.12 | — | — | — | |
| DoGLLM=GPT-4o, Training Strategy=Prompting with in-context learning2026.05 | 58.4 | 47.5 | — | — | — | — | — | — | — | — | — | |
| DualRLLM=Llama2-13B, Trainable Retriever GNN=true, Trainable Retriever LLM=true2026.06 | 58 | — | — | — | — | — | — | — | — | — | — | |
| RoGType=KG+LLM2026.03 | 57.8 | 56.2 | — | — | 77.3 | — | — | — | — | — | — | |
| SGRBackbone Model=ChatGPT2026.05 | 57.8 | — | 52.6 | — | — | — | — | — | — | — | — | |
| ToGLLM Backbone=ChatGPT2026.01 | 57.6 | — | — | — | — | — | — | — | — | — | — | |
| ToGBackbone=LLaMa2-70B2026.02 | 57.6 | — | — | — | — | — | — | — | — | — | — | |
| ToG+LLaMA2-70BType=KG+LLM2026.03 | 57.6 | 51.8 | — | — | 80.3 | — | — | — | — | — | — | |
| ToGLLM=Llama2-70B, Trainable Retriever LLM=true2026.06 | 57.6 | — | — | — | — | — | — | — | — | — | — |