Task Planning on HuggingGPT Human Evaluation Set 130 diverse requests (test)
0.9122Passing RateHuggingGPT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| HuggingGPTLLM=GPT-3.52023.03 | 0.9122 | 0.7847 | |
| HuggingGPTLLM=Vicuna-13b2023.03 | 0.7941 | 0.5841 | |
| HuggingGPTLLM=Alpaca-13b2023.03 | 0.5104 | 0.3217 |