Multi-turn Interaction on ALFWorld
82.8AccuracyQwen2.5-72B-CodeGym
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-72B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=72B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 82.8 | |
| Qwen2.5-32B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 80.8 | |
| Qwen2.5-72B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=72B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 80.4 | |
| Qwen2.5-14B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=14B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 72.8 | |
| Qwen2.5-32B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 66.8 | |
| QwQ-32B-CodeGymCoT Pattern=Long-CoT, Training Strategy=CodeGym, Model Series=QwQ, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 64.4 | |
| QwQ-32BCoT Pattern=Long-CoT, Training Strategy=Base, Model Series=QwQ, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 62.4 | |
| Qwen2.5-14B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=14B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 59.2 | |
| Qwen2.5-7B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=7B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 51.3 | |
| Qwen2.5-7B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=7B, T=0.7, top-p=0.95, Inference Averaging=5 runs, ReAct protocol=true2025.09 | 43.6 |