Context-faithful Multi-hop Reasoning on ConFiQA MR
45.4PcLLAMA2-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLAMA2-7BModel Version=chat2024.12 | 45.4 | 26.8 | 37.1 | 0.3 | |
| LLAMA2-13BModel Version=chat2024.12 | 43 | 33.7 | 43.9 | 0 | |
| LLAMA3-8BModel Version=instruct2024.12 | 30.6 | 44.1 | 59.1 | 0 | |
| MISTRAL-7BModel Version=instruct-v0.22024.12 | 21.7 | 37.9 | 63.5 | 0.2 | |
| QWEN2-7BModel Version=instruct2024.12 | 21.7 | 48.7 | 69.2 | 0 | |
| CHATGPT-42024.12 | 20.3 | 45.3 | 69.2 | 0.3 | |
| GEMINI-1.5-PRO2024.12 | 17.3 | 41.3 | 70.4 | 0 | |
| CHATGPT-4o2024.12 | 8.1 | 48.6 | 85.9 | 0 |