Text-based agent interaction on TextWorld Cooking (test)
75.5AccuracyDirect-Action (RL)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Direct-Action (RL)Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 75.5 | 31.9 | |
| Base ModelBackbone=Qwen3-4B-Instruct, Training Strategy=Inference-only2026.05 | 69.5 | 34.1 | |
| Agent-BRACEBackbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 69 | 44.6 | |
| Agent-BRACEBackbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 58.5 | 60.3 | |
| MEM1Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 52.5 | 48 | |
| Direct-Action (RL)Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 51.5 | 46.1 | |
| ReAct (RL)Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 34.5 | 44.4 | |
| PABUBackbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 33 | 73.1 | |
| PABUBackbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 32.5 | 75.6 | |
| ReActBackbone=Qwen2.5-3B-Instruct, Training Strategy=Inference-only2026.05 | 27.5 | 38.4 | |
| ReActBackbone=Qwen3-4B-Instruct, Training Strategy=Inference-only2026.05 | 13.5 | 24.4 | |
| ReAct (RL)Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 13 | 40.6 | |
| MEM1Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 10 | 10 | |
| Base ModelBackbone=Qwen2.5-3B-Instruct, Training Strategy=Inference-only2026.05 | 2.5 | 98.1 |