Multi-Hop Question Answering on Zh.QA
1.1Helmet Correctness ScoreExtAgents (Ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| ExtAgents (Ours)Base Model=gpt-4o-mini-2024-07-18, Input length (#tokens)=256k2025.05 | 1.1 | |
| Direct InputBase Model=gpt-4o-mini-2024-07-18, Input length (#tokens)=128k2025.05 | 1.04 | |
| LLM×MapReduceBase Model=gpt-4o-mini-2024-07-18, Input length (#tokens)=128k2025.05 | 1.04 | |
| Direct InputBase Model=Llama-3.1-8B-Instruct, Input length (#tokens)=128k2025.05 | 0.89 | |
| ExtAgents (Ours)Base Model=Llama-3.1-8B-Instruct, Input length (#tokens)=256k2025.05 | 0.85 | |
| LLM×MapReduceBase Model=Llama-3.1-8B-Instruct, Input length (#tokens)=256k2025.05 | 0.79 | |
| Direct InputBase Model=DeepSeek-R1-Distill-Llama-8B, Input length (#tokens)=32k2025.05 | 0.66 | |
| Chain of AgentsBase Model=Llama-3.1-8B-Instruct, Input length (#tokens)=32k2025.05 | 0.57 |