Agentic decision-making on ALFWorld (in-distribution)
90.5Success RateC-RF w/ NTF
Evaluation Results
| Method | Links | |
|---|---|---|
| C-RF w/ NTFModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 90.5 | |
| C-RF w/ NTFModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 87.4 | |
| GRPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 78.6 | |
| RLOOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 77.9 | |
| PPO (with critic)Model=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 77.1 | |
| GRPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 72.8 | |
| RLOOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 69.7 | |
| PPO (with critic)Model=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 54.4 | |
| ReflexionModel=Qwen2.5-7B-Instruct, Type=Prompting2026.06 | 42.7 | |
| ReActModel=Qwen2.5-7B-Instruct, Type=Prompting2026.06 | 31.2 | |
| ReflexionModel=Qwen2.5-1.5B-Instruct, Type=Prompting2026.06 | 21.8 | |
| Qwen2.5Model=Qwen2.5-7B-Instruct, Type=Prompting2026.06 | 14.8 | |
| ReActModel=Qwen2.5-1.5B-Instruct, Type=Prompting2026.06 | 12.8 | |
| Qwen2.5Model=Qwen2.5-1.5B-Instruct, Type=Prompting2026.06 | 4.1 |