Multi-hop Question Answering on HybridQA (dev)
71Answer CorrectnessA.DOT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| A.DOTInference Model=LLaMA-3 70B, Evaluation Judge=Mistral Large 2, Retrieval Architecture=SQL + Milvus2026.03 | 71 | 73 | |
| Standard RAGInference Model=LLaMA-3 70B, Evaluation Judge=Mistral Large 2, Retrieval Architecture=SQL + Milvus2026.03 | 56.2 | 62.3 | |
| ReActInference Model=LLaMA-3 70B, Evaluation Judge=Mistral Large 2, Retrieval Architecture=SQL + Milvus2026.03 | 40.2 | 44.3 | |
| LLM CompilerInference Model=LLaMA-3 70B, Evaluation Judge=Mistral Large 2, Retrieval Architecture=SQL + Milvus2026.03 | 27.8 | 30.8 |