Question Answering on StrategyQA (EM & Tool Calls)
80.1EMUALA-S+Oracle
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| UALA-S+OracleBackbone=LLaMA2-70B2024.01 | 80.1 | 298 | |
| RA-ISFModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 75.9 | — | |
| Iter-RetGenModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 72.3 | — | |
| UALA-M+BackoffBackbone=LLaMA2-70B2024.01 | 71.8 | 572 | |
| UALA-S+BackoffBackbone=LLaMA2-70B2024.01 | 71.6 | 298 | |
| SKRknnModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 70.1 | — | |
| UALA-SBackbone=LLaMA2-70B2024.01 | 69 | 298 | |
| Least-to-mostModel=GPT-3.5, Retrieval Configuration=Without Retrieval2024.03 | 68.5 | — | |
| IRCoTModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 67.9 | — | |
| Self-ConsistencyBackbone=LLaMA2-70B2024.01 | 67.7 | 0 | |
| UALA-S+OracleBackbone=ChatGPT2024.01 | 67.5 | 134 | |
| Self-RAG13BModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 67.2 | — | |
| UALA-M+BackoffBackbone=ChatGPT2024.01 | 66.9 | 234 | |
| ReAct+BackoffBackbone=LLaMA2-70B2024.01 | 66.8 | 890 | |
| RA-ISFModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 66.7 | — | |
| UALA-S+BackoffBackbone=ChatGPT2024.01 | 66.4 | 134 | |
| UALA-MBackbone=LLaMA2-70B2024.01 | 66.1 | 572 | |
| StandardBackbone=LLaMA2-70B2024.01 | 65.9 | 0 | |
| UALA-SBackbone=ChatGPT2024.01 | 65.5 | 134 | |
| DirectModel=GPT-3.5, Retrieval Configuration=Without Retrieval2024.03 | 65.2 | — | |
| RAGModel=GPT-3.5, Retrieval Configuration=With Retrieval2024.03 | 64.7 | — | |
| UALA-MBackbone=ChatGPT2024.01 | 63.8 | 234 | |
| CoTBackbone=LLaMA2-70B2024.01 | 63.8 | 0 | |
| REPLUGModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 62.9 | — | |
| ReAct+BackoffBackbone=ChatGPT2024.01 | 61.6 | 709 | |
| SKRknnModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 61.6 | — | |
| RAGModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 60.8 | — | |
| Least-to-mostModel=Llama-2-13b, Retrieval Configuration=Without Retrieval2024.03 | 60.5 | — | |
| IRCoTModel=Llama-2-13b, Retrieval Configuration=With Retrieval2024.03 | 59.1 | — | |
| Self-ConsistencyBackbone=ChatGPT2024.01 | 58.5 | 0 | |
| ReActBackbone=LLaMA2-70B2024.01 | 58.1 | 890 | |
| StandardBackbone=ChatGPT2024.01 | 57.6 | 0 | |
| CoTBackbone=ChatGPT2024.01 | 55.9 | 0 | |
| ReActBackbone=ChatGPT2024.01 | 55.5 | 709 | |
| Vanilla LMModel=Llama-2-13b, Retrieval Configuration=Without Retrieval2024.03 | 52.2 | — |