Text-based agent interaction on TextWorld Treasure (test)
81.5AccuracyAgent-BRACE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Agent-BRACEBackbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 81.5 | 32.1 | |
| Agent-BRACEBackbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 81 | 30 | |
| ReAct (RL)Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 74 | 16.5 | |
| PABUBackbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 73.5 | 37.2 | |
| PABUBackbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 72.5 | 34.4 | |
| Direct-Action (RL)Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 72.5 | 28 | |
| ReActBackbone=Qwen3-4B-Instruct, Training Strategy=Inference-only2026.05 | 69.5 | 10.3 | |
| Direct-Action (RL)Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 67.5 | 32.6 | |
| Base ModelBackbone=Qwen3-4B-Instruct, Training Strategy=Inference-only2026.05 | 65 | 30.3 | |
| MEM1Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 63.5 | 31.4 | |
| ReAct (RL)Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 55 | 32.7 | |
| ReActBackbone=Qwen2.5-3B-Instruct, Training Strategy=Inference-only2026.05 | 37 | 33.6 | |
| MEM1Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 30 | 47.7 | |
| Base ModelBackbone=Qwen2.5-3B-Instruct, Training Strategy=Inference-only2026.05 | 7.5 | 93.2 |