Question Answering on ASQA (in-domain)
47.21EMIn-context Prompting (gpt-4o)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| In-context Prompting (gpt-4o)Type=Training-Free, Model=gpt-4o2025.03 | 47.21 | 75.11 | 71.52 | 18.57 | |
| GenDiE (Hierarchical inference)Type=Ours, Decoding=Hierarchical inference (beam3-beam3)2025.03 | 45.75 | 84.9 | 82.03 | 21.52 | |
| Standard SFT (beam3)Type=Training-based, Model=llama3.1-8b, Decoding=beam32025.03 | 44.64 | 73.88 | 66.47 | 19.51 | |
| GenDiE (Vanilla beam search)Type=Ours, Decoding=Vanilla beam search (beam3)2025.03 | 43.71 | 82.3 | 79.42 | 18.88 | |
| GenDiE (Greedy search)Type=Ours, Decoding=Greedy search2025.03 | 43.61 | 73.87 | 69.16 | 18.57 | |
| GenDiE_answer-levelType=Training-based, Decoding=greedy2025.03 | 43.13 | 64.48 | 56.51 | 19.2 | |
| GenDiE_gold-answerType=Training-based, Decoding=greedy2025.03 | 42.88 | 71.66 | 64.76 | 17.62 | |
| CADType=Training-Free, Model=llama3.1-8b, Protocol=Standard SFT, Decoding=greedy2025.03 | 42.72 | 72.07 | 66.18 | 17.61 | |
| Standard SFT (greedy)Type=Training-based, Model=llama3.1-8b, Decoding=greedy2025.03 | 41.93 | 65.27 | 57.54 | 16.35 | |
| Extractive Sentence Selection (stella-1.5b)Type=Training-Free, Model=stella-1.5b2025.03 | 41 | — | — | 17.19 | |
| In-context Prompting (llama3.1-8b)Type=Training-Free, Model=llama3.1-8b2025.03 | 40.3 | 76.9 | 70.39 | 14.35 | |
| Extractive Sentence Selection (instructor-large)Type=Training-Free, Model=instructor-large2025.03 | 36.58 | — | — | 14.45 |