Contextual Extraction on NQ-Open, TriviaQA, and HotpotQA (1k samples/corpus, test)
62Accuracy (Zero-Shot)GPT-4o-mini
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| GPT-4o-miniModel Name=GPT-4o-mini2026.04 | 62 | 90.1 | 89.2 | 47.1 | 0.777 | -7,779.254 | 2.982 | |
| Qwen 72BModel Size=72B2026.04 | 55.9 | 91.8 | 91.7 | 42.5 | 0.868 | -20.882 | 0.389 | |
| Qwen 1.5BModel Size=1.5B2026.04 | 53.6 | 95.9 | 95.5 | 34.9 | 0.864 | 0.62 | 0.041 | |
| Qwen 7BModel Size=7B2026.04 | 51.9 | 92.3 | 92.3 | 33.1 | 0.864 | -74.128 | 0.093 |