Long-Context QA on LongMemEval
11.6AccuracyC-DIC
Evaluation Results
| Method | Links | |
|---|---|---|
| C-DICBackbone=Llama-2-Chat-7B, Zero-shot=true, Automatic judge=GPT-4o2026.06 | 11.6 | |
| Full promptingBackbone=Llama-2-Chat-7B, Zero-shot=true, Automatic judge=GPT-4o2026.06 | 8.6 | |
| ICAE (incremental)Backbone=Llama-2-Chat-7B, Zero-shot=true, Automatic judge=GPT-4o2026.06 | 1 | |
| ICAE (one-shot)Backbone=Llama-2-Chat-7B, Zero-shot=true, Automatic judge=GPT-4o2026.06 | 0.4 |