Question Answering on HotpotQA (Uncertainty Evaluation)
66.8AccuracyIterRetGen
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| IterRetGen2025.03 | 66.8 | — | — | — | — | |
| L-RAGRetrieval Count (K)=82025.03 | 64.1 | — | — | — | — | |
| IRCoT2025.03 | 61.5 | — | — | — | — | |
| VanillaRAGRetrieval Count (K)=82025.03 | 59.8 | — | — | — | — | |
| VanillaRAGRetrieval Count (K)=42025.03 | 59 | — | — | — | — | |
| L-RAGRetrieval Count (K)=42025.03 | 55.5 | — | — | — | — | |
| SelfAsk2025.03 | 37.3 | — | — | — | — | |
| GPT-4Prompt Type=vanilla prompt2024.02 | 31.98 | 0.3437 | 0.0561 | 0.3926 | 0.5513 | |
| HyDE2025.03 | 24.8 | — | — | — | — | |
| GPT-InstructPrompt Type=vanilla prompt2024.02 | 23.3 | 0.2188 | 0.0144 | 0.5626 | 0.423 | |
| No RetrievalRetrieval=None2025.03 | 23 | — | — | — | — | |
| ChatGPTPrompt Type=vanilla prompt2024.02 | 19.51 | 0.5679 | 0.0376 | 0.2747 | 0.6877 | |
| VicunaPrompt Type=vanilla prompt2024.02 | 14.47 | 0.0571 | 0.003 | 0.8012 | 0.1957 | |
| LLaMA2Prompt Type=vanilla prompt2024.02 | 11.68 | 0.4484 | 0.023 | 0.456 | 0.5209 |