Question Answering on HotpotQA (LLM Accuracy and EM)
36.3Exact Match (EM)LUMOS-IQA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LUMOS-IQAAgent Model=LLAMA-2-13B, QA Tool=GPT-4, fine-tuning=true2023.11 | 36.3 | 57.4 | |
| LUMOS-IQAAgent Model=LLAMA-2-7B, QA Tool=GPT-4, fine-tuning=true2023.11 | 36 | 56.8 | |
| ReAcTAgent Model=GPT-3.5-turbo, QA Tool=GPT-3.5-turbo, fine-tuning=false2023.11 | 32.4 | 40.8 | |
| LUMOS-IQAAgent Model=LLAMA-2-13B, QA Tool=GPT-3.5-turbo, fine-tuning=true2023.11 | 31.4 | 50.2 | |
| ReWOOAgent Model=GPT-3.5-turbo, QA Tool=GPT-3.5-turbo, fine-tuning=false2023.11 | 30.4 | 42.4 | |
| LUMOS-IQAAgent Model=LLAMA-2-7B, QA Tool=GPT-3.5-turbo, fine-tuning=true2023.11 | 29.4 | 45.9 | |
| FiReActAgent Model=CodeLLAMA-34B, QA Tool=CodeLLAMA-34B, fine-tuning=true2023.11 | 27.8 | — | |
| FiReActAgent Model=LLAMA-2-7B, QA Tool=LLAMA-2-7B, fine-tuning=true2023.11 | 26.2 | — | |
| LUMOS-OQAAgent Model=LLAMA-2-7B, QA Tool=GPT-3.5-turbo, fine-tuning=true2023.11 | 24.9 | 39.2 | |
| LUMOS-IQAAgent Model=LLAMA-2-7B, QA Tool=LLAMA-2-7B, fine-tuning=true2023.11 | 23.5 | 37.3 | |
| GPT-3.5-CoTAgent Model=GPT-3.5-turbo, QA Tool=GPT-3.5-turbo, fine-tuning=false2023.11 | 22.4 | 37.8 | |
| AgentLMAgent Model=LLAMA-2-7B, QA Tool=LLAMA-2-7B, fine-tuning=false2023.11 | 22.3 | — | |
| ReWOO-openAgent Model=LLAMA-7B, QA Tool=GPT-3.5-turbo, fine-tuning=true2023.11 | — | 37 |