Question Answering on WebQSP
96.7Hit@1Paths-over-Graph
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Paths-over-GraphSetting=LLM w/ LLM Retriever, LLM=GPT-42026.06 | 96.7 | — | — | — | |
| DoGSetting=LLM w/ LLM Retriever, LLM=GPT-42026.06 | 91 | — | — | — | |
| Plan-on-GraphSetting=LLM w/ LLM Retriever, LLM=GPT-42026.06 | 87.3 | — | — | — | |
| AGE AMARSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama2-7B2026.06 | 86.5 | — | — | — | |
| EMBRAGType=LLMs+KGs, Backbone=LLaMA2-Chat-7B2026.02 | 86.4 | — | — | — | |
| AGE AMARSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama2-13B2026.06 | 86.2 | — | — | — | |
| ReKnoSSetting=LLM w/ LLM Retriever, LLM=GPT-42026.06 | 84.9 | — | — | — | |
| AMARSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama2-7B2026.06 | 84.3 | — | — | — | |
| AMARSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama2-13B2026.06 | 83.3 | — | — | — | |
| KG-AgentSetting=LLM w/ LLM Retriever, LLM=Llama2-7B2026.06 | 83.3 | — | — | — | |
| ToGBackbone=GPT-4, Reasoning Strategy=ToG2025.12 | 82.6 | — | — | — | |
| SGRBackbone=GPT-4, Reasoning Strategy=SGR2025.12 | 82.6 | 80.8 | — | — | |
| ToGSetting=LLM w/ LLM Retriever, LLM=GPT-42026.06 | 82.6 | — | — | — | |
| Prior FT SOTAEvaluation Protocol=Fine-tuned2025.12 | 82.1 | — | — | — | |
| AGE G-RetrieverSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama3.1-8B2026.06 | 80.3 | — | — | — | |
| SGRBackbone=ChatGPT, Reasoning Strategy=SGR2025.12 | 80.1 | 78.4 | — | — | |
| AGE G-RetrieverSetting=Frozen LLM w/ Graph Embedding, LLM=Llama3.1-8B2026.06 | 78.3 | — | — | — | |
| AGE G-RetrieverSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama3.2-3B2026.06 | 77.3 | — | — | — | |
| Qwen3-235BFramework=StructGPT2026.04 | 76.49 | — | — | 22.84 | |
| ToGBackbone=ChatGPT, Reasoning Strategy=ToG2025.12 | 76.2 | — | — | — | |
| MixedFramework=StructGPT2026.04 | 75.8 | — | — | 11.79 | |
| ChatGPT+CoTType=LLMs2026.02 | 75.6 | — | — | — | |
| Qwen3-235BFramework=Readi2026.04 | 75.17 | — | — | 58.45 | |
| SGRBackbone=Cypher LLM, Reasoning Strategy=SGR2025.12 | 74.5 | 70.6 | — | — | |
| Prior Prompting SOTAEvaluation Protocol=Prompting2025.12 | 74.4 | — | — | — | |
| MixedFramework=Readi2026.04 | 74.31 | — | — | 25.97 | |
| Qwen3-235BFramework=ToG2026.04 | 74.27 | — | — | 64.25 | |
| G-Retriever w/ LoRASetting=Tuned LLM2024.02 | 73.79 | — | — | — | |
| AGE G-RetrieverSetting=Frozen LLM w/ Graph Embedding, LLM=Llama3.2-3B2026.06 | 73.5 | — | — | — | |
| MixedFramework=ToG2026.04 | 73.2 | — | — | 26.77 | |
| StructGPTBackbone=ChatGPT, Reasoning Strategy=StructGPT2025.12 | 72.6 | — | — | — | |
| Qwen2.5-7BFramework=StructGPT2026.04 | 71.76 | — | — | — | |
| TransferNetType=Embedding2026.02 | 71.4 | — | — | — | |
| G-RetrieverSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama3.2-3B2026.06 | 71.4 | — | — | — | |
| G-RetrieverSetting=Frozen LLM w/ Graph Embedding, LLM=Llama3.2-3B2026.06 | 71.3 | — | — | — | |
| G-RetrieverSetting=Frozen LLM w/ PT2024.02 | 70.49 | — | — | — | |
| G-RetrieverSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama2-7B2026.06 | 70.2 | — | — | — | |
| Qwen2.5-7BFramework=Readi2026.04 | 69.83 | — | — | — | |
| AGE G-RetrieverSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama3.2-1B2026.06 | 69.1 | — | — | — | |
| PullNetType=Retrieval2026.02 | 68.1 | — | — | — | |
| G-RetrieverSetting=Frozen LLM w/ Graph Embedding, LLM=Llama2-7B2026.06 | 68.1 | — | — | — | |
| BiNetType=Embedding2026.02 | 67.6 | — | — | — | |
| Qwen2.5-7BFramework=ToG2026.04 | 67.25 | — | — | — | |
| ChatGPTType=LLMs2026.02 | 66.8 | — | — | — | |
| EmbedKGQAType=Embedding2026.02 | 66.6 | — | — | — | |
| GraftNetType=Retrieval2026.02 | 66.4 | — | — | — | |
| LoRASetting=Tuned LLM2024.02 | 66.03 | — | — | — | |
| G-RetrieverSetting=Frozen LLM w/ Graph Embedding + PEFT(LoRA), LLM=Llama3.2-1B2026.06 | 65.3 | — | — | — | |
| IO PromptBackbone=ChatGPT, Reasoning Strategy=IO Prompt2025.12 | 63.3 | 58.2 | — | — | |
| AGE G-RetrieverSetting=Frozen LLM w/ Graph Embedding, LLM=Llama3.2-1B2026.06 | 62.5 | — | — | — | |
| CoTBackbone=ChatGPT, Reasoning Strategy=CoT2025.12 | 62.2 | 57.7 | — | — | |
| G-RetrieverSetting=Frozen LLM w/ Graph Embedding, LLM=Llama3.2-1B2026.06 | 60.1 | — | — | — | |
| GraphTokenSetting=Frozen LLM w/ PT2024.02 | 57.05 | — | — | — | |
| KAPINGSetting=Inference-only2024.02 | 52.64 | — | — | — | |
| Zero-CoTSetting=Inference-only2024.02 | 51.3 | — | — | — | |
| Prompt tuningSetting=Frozen LLM w/ PT2024.02 | 48.34 | — | — | — | |
| KV-MemType=Embedding2026.02 | 46.7 | — | — | — | |
| Zero-shotSetting=Inference-only2024.02 | 41.06 | — | — | — | |
| COT-BAGSetting=Inference-only2024.02 | 39.6 | — | — | — | |
| CureLLM2026.06 | — | 72.58 | — | — | |
| G-retrieverGenerator=Qwen2-7B2025.10 | — | 25.74 | 35.45 | — | |
| G-retrieverGenerator=Llama3-8B2025.10 | — | 22.67 | 32.26 | — | |
| G-retrieverGenerator=Finetuned-7B2025.10 | — | 30.34 | 43.49 | — | |
| G-Retriever2026.06 | — | 70.49 | — | — | |
| Graph-S^3Generator=Qwen2-7B2025.10 | — | 36.24 | 47.88 | — | |
| Graph-S^3Generator=Llama3-8B2025.10 | — | 32.31 | 43.26 | — | |
| Graph-S^3Generator=Finetuned-7B2025.10 | — | 44.29 | 58.45 | — | |
| GraphToken2026.06 | — | 57.05 | — | — | |
| KG-AgentGenerator=Qwen2-7B2025.10 | — | 29.66 | 38.02 | — | |
| KG-AgentGenerator=Llama3-8B2025.10 | — | 32.04 | 41.99 | — | |
| KG-AgentGenerator=Finetuned-7B2025.10 | — | 42.6 | 55.38 | — | |
| LightRAGGenerator=Qwen2-7B2025.10 | — | 18.39 | 31.67 | — | |
| LightRAGGenerator=Llama3-8B2025.10 | — | 15.85 | 36.66 | — | |
| LightRAGGenerator=Finetuned-7B2025.10 | — | 17.38 | 32.59 | — | |
| MuseGraph2026.06 | — | 70.61 | — | — | |
| No graphGenerator=Qwen2-7B2025.10 | — | 5.16 | 8.11 | — | |
| No graphGenerator=Llama3-8B2025.10 | — | 8.97 | 15.69 | — | |
| No graphGenerator=Finetuned-7B2025.10 | — | 9.21 | 14.88 | — | |
| No retrieverGenerator=Qwen2-7B2025.10 | — | 0.25 | 1.83 | — | |
| No retrieverGenerator=Llama3-8B2025.10 | — | 0.18 | 1.97 | — | |
| No retrieverGenerator=Finetuned-7B2025.10 | — | 0.37 | 2.26 | — | |
| Prompt tuning2026.06 | — | 48.34 | — | — | |
| RAG/1hopGenerator=Qwen2-7B2025.10 | — | 27.89 | 38.57 | — | |
| RAG/1hopGenerator=Llama3-8B2025.10 | — | 24.82 | 35.28 | — | |
| RAG/1hopGenerator=Finetuned-7B2025.10 | — | 28.87 | 41.48 | — | |
| RAG/2hopGenerator=Qwen2-7B2025.10 | — | 14.07 | 24.47 | — | |
| RAG/2hopGenerator=Llama3-8B2025.10 | — | 11.06 | 22.94 | — | |
| RAG/2hopGenerator=Finetuned-7B2025.10 | — | 14.93 | 27.76 | — | |
| RAG/3hopGenerator=Qwen2-7B2025.10 | — | 1.54 | 7.94 | — | |
| RAG/3hopGenerator=Llama3-8B2025.10 | — | 1.04 | 6.67 | — | |
| RAG/3hopGenerator=Finetuned-7B2025.10 | — | 1.54 | 7.81 | — | |
| ToGGenerator=Qwen2-7B2025.10 | — | 6.14 | 9.79 | — | |
| ToGGenerator=Llama3-8B2025.10 | — | 8.85 | 14.28 | — | |
| ToGGenerator=Finetuned-7B2025.10 | — | 5.04 | 9.43 | — |