Document-grounded Question Answering on Long-Context Noise Dataset (800 samples) (test)
296ARCIP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CIPBackbone Model=GPT-4o, Prompting Strategy=Causal Prompting, Latency (s)=3.832025.12 | 296 | 40 | |
| RAG EnhancementBackbone Model=GPT-4o, Prompting Strategy=Retrieval-augmented generation (top-5), Latency (s)=8.342025.12 | 51 | 12 | |
| CoTBackbone Model=GPT-4o, Prompting Strategy=Chain-of-Thought, Latency (s)=7.152025.12 | 42 | 8 | |
| Direct PromptingBackbone Model=GPT-4o, Prompting Strategy=Standard query + context without any enhancement, Latency (s)=6.922025.12 | 35 | 5 |