Question Answering on ConFiQA (out-of-domain)
84.63HitGenDiE (Hierarchical inference)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GenDiE (Hierarchical inference)Type=Ours, Decoding=Hierarchical inference (beam3-beam3)2025.03 | 84.63 | 80.73 | 80.69 | |
| GenDiE (Vanilla beam search)Type=Ours, Decoding=Vanilla beam search (beam3)2025.03 | 80.95 | 78.54 | 69.33 | |
| GenDiE_answer-levelType=Training-based, Decoding=greedy2025.03 | 77.91 | 69.81 | 66.17 | |
| GenDiE (Greedy search)Type=Ours, Decoding=Greedy search2025.03 | 73.72 | 72.32 | 70.17 | |
| GenDiE_gold-answerType=Training-based, Decoding=greedy2025.03 | 65.27 | 64.72 | 59.89 | |
| Standard SFT (beam3)Type=Training-based, Model=llama3.1-8b, Decoding=beam32025.03 | 59.68 | 66.03 | 62.9 | |
| In-context Prompting (llama3.1-8b)Type=Training-Free, Model=llama3.1-8b2025.03 | 55.14 | 50.67 | 30.78 | |
| CADType=Training-Free, Model=llama3.1-8b, Protocol=Standard SFT, Decoding=greedy2025.03 | 49.38 | 77.46 | 52.79 | |
| Extractive Sentence Selection (stella-1.5b)Type=Training-Free, Model=stella-1.5b2025.03 | 42.31 | — | — | |
| Extractive Sentence Selection (instructor-large)Type=Training-Free, Model=instructor-large2025.03 | 39.17 | — | — | |
| In-context Prompting (gpt-4o)Type=Training-Free, Model=gpt-4o2025.03 | 35.96 | 44.56 | 34.11 | |
| Standard SFT (greedy)Type=Training-based, Model=llama3.1-8b, Decoding=greedy2025.03 | 34.58 | 44.94 | 40.74 |