Multi-doc QA on HotpotQA
45AccuracyDense
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DenseModel=QwQ 32B2025.07 | 45 | 0 | |
| DenseModel=Phi 4 reasoning plus2025.07 | 43 | 0 | |
| ReasonCacheModel=QwQ 32B2025.07 | 41 | 16.12 | |
| ReasonCacheModel=Phi 4 reasoning plus2025.07 | 40 | 16.2 | |
| DenseModel=DeepSeek R1 Distill Qwen 32B2025.07 | 37.5 | 0 | |
| ReasonCacheModel=DeepSeek R1 Distill Qwen 32B2025.07 | 36 | 16.51 | |
| QuestModel=QwQ 32B2025.07 | 36 | 15.82 | |
| SnapKVModel=QwQ 32B2025.07 | 32 | 15.7 | |
| SnapKVModel=Phi 4 reasoning plus2025.07 | 29.2 | 15.5 | |
| QuestModel=Phi 4 reasoning plus2025.07 | 29 | 15.75 | |
| SnapKVModel=DeepSeek R1 Distill Qwen 32B2025.07 | 28.5 | 16.22 | |
| StreamingLLMModel=QwQ 32B2025.07 | 28.2 | 16.51 | |
| StreamingLLMModel=Phi 4 reasoning plus2025.07 | 27.8 | 14.8 | |
| QuestModel=DeepSeek R1 Distill Qwen 32B2025.07 | 24 | 16.02 | |
| StreamingLLMModel=DeepSeek R1 Distill Qwen 32B2025.07 | 23.5 | 15.8 |