Claim Verification on FactKG (test)
86.8Average AccuracySimGRAG
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| SimGRAGTraining Paradigm=KG-driven RAG without training, Backbone=Llama 3 70B2024.12 | 86.8 | — | — | — | — | — | — | |
| ClaimPKGSpecialized LLM=Llama-3B*, General LLM=Qwen-72B, Evidence Retrieval=True2025.05 | 85.22 | 100 | 85.27 | 86.9 | 84.02 | 78.71 | 91.2 | |
| ClaimPKGSpecialized LLM=Llama-3B*, General LLM=Llama-70B, Evidence Retrieval=True2025.05 | 84.64 | 100 | 84.58 | 84.2 | 85.68 | 78.49 | 90.26 | |
| ClaimPKGAblation=w/o Trie Constraint, Specialized LLM=Llama-3B*, General LLM=Llama-70B, Evidence Retrieval=True2025.05 | 82.74 | 87.5 | 82.5 | 83.24 | 83.82 | 76.13 | 88.01 | |
| ClaimPKGSpecialized LLM=Llama-3B*, General LLM=GPT-4o-mini, Evidence Retrieval=True2025.05 | 81.05 | 100 | 85.1 | 72.64 | 84.23 | 72.26 | 91.01 | |
| GEARTraining Paradigm=Supervised task-specific2024.12 | 77.7 | — | — | — | — | — | — | |
| ClaimPKGAblation=Few-shot Specialized LLM, Specialized LLM=Llama-70B, General LLM=Llama-70B, Evidence Retrieval=True2025.05 | 77.63 | 86.52 | 77.99 | 81.89 | 77.8 | 68.82 | 81.65 | |
| GEARBackbone=Finetuned BERT, Evidence Retrieval=True2025.05 | 76.65 | — | 79.72 | 79.19 | 78.63 | 68.39 | 77.34 | |
| KAPINGTraining Paradigm=KG-driven RAG without training, Backbone=Llama 3 70B2024.12 | 75.5 | — | — | — | — | — | — | |
| KG-GPTBackbone=Llama-70B, Strategy=Few-shot, Evidence Retrieval=True2025.05 | 74.7 | — | 70.91 | 65.06 | 86.64 | 58.87 | 92.02 | |
| KELPTraining Paradigm=KG-driven RAG with training, Backbone=Llama 3 70B2024.12 | 73.3 | — | — | — | — | — | — | |
| KG-GPTBackbone=Qwen-72B, Strategy=Few-shot, Evidence Retrieval=True2025.05 | 73.12 | — | 67.31 | 60.08 | 89.14 | 58.19 | 90.87 | |
| KG-GPTTraining Paradigm=KG-driven RAG without training, Backbone=Llama 3 70B, Oracle Entities=true2024.12 | 69.5 | — | — | — | — | — | — | |
| Llama-70BStrategy=Zero-shot CoT, Evidence Retrieval=False2025.05 | 69.07 | — | 64.34 | 64.62 | 72.47 | 65.58 | 78.32 | |
| ChatGPTTraining Paradigm=Pre-trained LLM, Prompting=12-shots2024.12 | 68.5 | — | — | — | — | — | — | |
| Llama 3 70BTraining Paradigm=Pre-trained LLM, Prompting=12-shots2024.12 | 68.4 | — | — | — | — | — | — | |
| Qwen-72BStrategy=Zero-shot CoT, Evidence Retrieval=False2025.05 | 67.49 | — | 62.91 | 62.2 | 74.04 | 62.32 | 75.98 | |
| ClaimPKGAblation=w/o Incomplete Retrieval, Specialized LLM=Llama-3B*, General LLM=Llama-70B, Evidence Retrieval=True2025.05 | 65.08 | 100 | 68.8 | 51.25 | 67.84 | 61.29 | 76.22 | |
| GPT-4o-miniStrategy=Zero-shot CoT, Evidence Retrieval=False2025.05 | 64.51 | — | 61.91 | 59.45 | 69.51 | 60.87 | 70.83 | |
| G-RetrieverTraining Paradigm=KG-driven RAG with training, Backbone=Llama 3 70B, Oracle Entities=true2024.12 | 61.4 | — | — | — | — | — | — |