Interactive environment task success on ALFWorld (test)
91.79Overall Success RateAdaPlanner
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| AdaPlannerMethod Category=Explicit Closed-Loop Methods with Plan Refinement, Backbone LLM=GPT-3 (text-davinci-002)2023.05 | 91.79 | 100 | 96.77 | 95.65 | 100 | — | 47.06 | 100 | |
| AUTOGUIDE + ReflexionOffline data=true, Context aware=true, Base agent model=GPT-3.5-turbo, Reflexion=true, Reflexion model=GPT-4-turbo2024.03 | 88.1 | — | — | — | — | — | — | — | |
| ReflexionMethod Category=Implicit Closed-Loop Methods with Fixed Plan, Backbone LLM=GPT-3+3.5 (text-davinci-002 and text-davinci-003)2023.05 | 88.06 | 75 | 90.32 | 91.3 | 90.48 | — | 94.12 | 88.89 | |
| AdaPlannerMethod Category=Explicit Closed-Loop Methods with Plan Refinement, Backbone LLM=GPT-3.5 (gpt-3.5-turbo)2023.05 | 80.6 | 77.78 | 93.55 | 69.57 | 93.65 | — | 78.43 | 62.96 | |
| AUTOGUIDEOffline data=true, Context aware=true, Base agent model=GPT-3.5-turbo, Reflexion=false2024.03 | 79.1 | — | — | — | — | — | — | — | |
| ExpeL + ReflexionOffline data=true, Context aware=false, Base agent model=GPT-3.5-turbo, Reflexion=true, Reflexion model=GPT-4-turbo2024.03 | 71.6 | — | — | — | — | — | — | — | |
| ReActSelection criteria=best of 6, Decoding=greedy2022.10 | 71 | 92 | 58 | 96 | 86 | 78 | 41 | — | |
| ReAct + ReflexionOffline data=false, Context aware=false, Base agent model=GPT-3.5-turbo, Reflexion=true, Reflexion model=GPT-4-turbo2024.03 | 67.2 | — | — | — | — | — | — | — | |
| ReActMethod Category=Implicit Closed-Loop Methods with Fixed Plan, Backbone LLM=GPT-3 (text-davinci-002)2023.05 | 61.94 | 66.67 | 41.94 | 91.03 | 80.95 | — | 35.29 | 55.56 | |
| ExpeLOffline data=true, Context aware=false, Base agent model=GPT-3.5-turbo, Reflexion=false2024.03 | 59 | — | — | — | — | — | — | — | |
| ReActSelection criteria=average, Decoding=greedy2022.10 | 57 | 65 | 39 | 83 | 76 | 55 | 24 | — | |
| ReActOffline data=false, Context aware=false, Base agent model=GPT-3.5-turbo, Reflexion=false2024.03 | 54.5 | — | — | — | — | — | — | — | |
| ReAct-IMSelection criteria=best of 6, Decoding=greedy, Prompting=IM-like dense external feedback2022.10 | 53 | 62 | 68 | 87 | 57 | 39 | 33 | — | |
| ReflexionMethod Category=Implicit Closed-Loop Methods with Fixed Plan, Backbone LLM=GPT-3.5 (gpt-3.5-turbo)2023.05 | 52.99 | 50 | 41.94 | 65.22 | 52.38 | — | 47.06 | 66.67 | |
| ReAct-IMSelection criteria=average, Decoding=greedy, Prompting=IM-like dense external feedback2022.10 | 48 | 55 | 59 | 60 | 55 | 23 | 24 | — | |
| ReActMethod Category=Implicit Closed-Loop Methods with Fixed Plan, Backbone LLM=GPT-3.5 (gpt-3.5-turbo)2023.05 | 47.76 | 37.5 | 64.52 | 69.57 | 42.86 | — | 17.65 | 38.89 | |
| ActSelection criteria=best of 6, Decoding=greedy2022.10 | 45 | 88 | 42 | 74 | 67 | 72 | 41 | — | |
| BUTLERSelection criteria=best of 8, Decoding=beam search2022.10 | 37 | 46 | 39 | 74 | 100 | 22 | 24 | — | |
| BUTLERMethod Category=Training-Based Methods2023.05 | 37 | 46 | 39 | 74 | 100 | — | 24 | 22 | |
| BUTLER_gSelection criteria=best of 8, Decoding=greedy2022.10 | 22 | 33 | 26 | 70 | 76 | 17 | 12 | — |