Knowledge-intensive Dialogue on HybriDialogue
85.89Factuality ScoreFine-Refine
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Fine-RefineBackbone=Llama-3.1-8B-Instruct, Refinement Iterations=32026.02 | 85.89 | 41.62 | 2.734 | 2.552 | |
| Self-RefineBackbone=Llama-3.1-8B-Instruct, Refinement Iterations=32026.02 | 81.73 | 61.18 | 2.883 | 2.598 | |
| Fine-RefineBackbone=Qwen3-8B, Refinement Iterations=32026.02 | 80.85 | 48.37 | 2.81 | 2.835 | |
| Fine-RefineBackbone=Mistral-7B-Instruct-v0.3, Refinement Iterations=32026.02 | 80.62 | 45.97 | 2.879 | 2.775 | |
| Self-RAG2026.02 | 79.74 | 46.39 | 2.162 | 2.238 | |
| Llama-3.1-8B-InstructBackbone=Llama-3.1-8B-Instruct2026.02 | 79.74 | 50.56 | 2.807 | 2.724 | |
| Self-RefineBackbone=Mistral-7B-Instruct-v0.3, Refinement Iterations=32026.02 | 79.65 | 63.1 | 2.894 | 2.755 | |
| Mistral-7B-Instruct-v0.3Backbone=Mistral-7B-Instruct-v0.32026.02 | 77.41 | 50.28 | 2.935 | 2.948 | |
| Self-RefineBackbone=Qwen3-8B, Refinement Iterations=32026.02 | 75.66 | 61.99 | 2.875 | 2.868 | |
| Qwen3-8BBackbone=Qwen3-8B2026.02 | 73.22 | 54.26 | 2.878 | 2.933 | |
| G-Retriever2026.02 | 70.29 | 55.77 | 2.824 | 2.788 |