Knowledge Base Question Answering on GrailQA (test)
91.76F1Pangu (T5-base)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Pangu (T5-base)Learning Mode=Bottom-up Parser, Data=full train data, Backbone=T5-base, Zero-shot=true, Entity Linker=Golden2024.06 | 91.76 | 88.3 | |
| ArcaneQALearning Mode=Bottom-up Parser, Data=full train data, Zero-shot=true, Entity Linker=Golden2024.06 | 81.81 | 78.52 | |
| DARA (w/Llama-2-13B)Learning Mode=Fine-tuned LLM Agent, Backbone=Llama-2-13B, Zero-shot=true, Entity Linker=Golden2024.06 | 80.35 | 77.03 | |
| DARA (w/Mistral-7B)Learning Mode=Fine-tuned LLM Agent, Backbone=Mistral-7B, Zero-shot=true, Entity Linker=Golden2024.06 | 80.16 | 76.88 | |
| DARA (w/CodeLlama-7B)Learning Mode=Fine-tuned LLM Agent, Backbone=CodeLlama-7B, Zero-shot=true, Entity Linker=Golden2024.06 | 78.61 | 75.26 | |
| DARA (w/Llama-2-7B)Learning Mode=Fine-tuned LLM Agent, Backbone=Llama-2-7B, Zero-shot=true, Entity Linker=Golden2024.06 | 77.71 | 75.05 | |
| FuSIC-KBQAEvaluation Protocol=Few-shot, Category=Fusion, Retriever=TIARA and Pangu2023.11 | 69.1 | — | |
| FuSIC-KBQAEvaluation Protocol=Zero-shot, Category=Fusion, Retriever=TIARA and Pangu2023.11 | 67.5 | — | |
| AgentBench (GPT-4)Learning Mode=Off-the-shelf (in-context learning), Backbone=GPT-4, Zero-shot=true, Entity Linker=Golden2024.06 | 65.89 | 63.56 | |
| FuSIC-KBQAEvaluation Protocol=Few-shot, Category=Fusion, Retriever=Pangu2023.11 | 64.5 | — | |
| FuSIC-KBQAEvaluation Protocol=Zero-shot, Category=Fusion, Retriever=Pangu2023.11 | 63.8 | — | |
| PanguEvaluation Protocol=Few-shot, Category=Supervised2023.11 | 63.2 | — | |
| PanguEvaluation Protocol=Zero-shot, Category=Supervised2023.11 | 60.9 | — | |
| FuSIC-KBQAEvaluation Protocol=Few-shot, Category=Fusion, Retriever=TIARA2023.11 | 60.3 | — | |
| AgentBench-7BLearning Mode=Fine-tuned LLM Agent, Backbone=7B, Zero-shot=true, Entity Linker=Golden2024.06 | 59.28 | 56.96 | |
| FuSIC-KBQAEvaluation Protocol=Zero-shot, Category=Fusion, Retriever=TIARA2023.11 | 59.2 | — | |
| AgentLM-13BLearning Mode=Fine-tuned LLM Agent, Backbone=13B, Zero-shot=true, Entity Linker=Golden2024.06 | 55.01 | 52.72 | |
| gf-LLMEvaluation Protocol=Zero-shot, Category=LLM ICL2023.11 | 53.4 | — | |
| TIARAEvaluation Protocol=Zero-shot, Category=Supervised2023.11 | 52.8 | — | |
| gf-LLMEvaluation Protocol=Few-shot, Category=LLM ICL2023.11 | 52.5 | — | |
| TIARAEvaluation Protocol=Few-shot, Category=Supervised2023.11 | 50.1 | — | |
| KB-BinderEvaluation Protocol=Zero-shot, Category=LLM ICL2023.11 | 45 | — | |
| BERT-RankingEvaluation Protocol=Few-shot, Category=LLM ICL2023.11 | 44.1 | — | |
| BERT-RankingEvaluation Protocol=Zero-shot, Category=LLM ICL2023.11 | 42.4 | — | |
| AgentBench (Llama-2-chat-70B)Learning Mode=Off-the-shelf (in-context learning), Backbone=Llama-2-chat-70B, Zero-shot=true, Entity Linker=Golden2024.06 | 35.72 | 33.2 | |
| KB-BinderEvaluation Protocol=Few-shot, Category=LLM ICL2023.11 | 35.2 | — | |
| AgentLM-7BLearning Mode=Fine-tuned LLM Agent, Backbone=7B, Zero-shot=true, Entity Linker=Golden2024.06 | 15.27 | 14.45 | |
| Pangu (T5-Large)Learning Mode=Bottom-up Parser, Data=full train data, Backbone=T5-Large, Zero-shot=true, Entity Linker=Golden2024.06 | — | 67.21 |