Text-based agent interaction on TextWorld Quest (test)
88AccuracyAgent-BRACE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Agent-BRACEBackbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 88 | 30.5 | |
| PABUBackbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 82.2 | 29.1 | |
| Agent-BRACEBackbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 78.5 | 37.3 | |
| ReAct (RL)Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 75.5 | 18.2 | |
| Direct-Action (RL)Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 74 | 29.6 | |
| PABUBackbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 73 | 37 | |
| Base ModelBackbone=Qwen3-4B-Instruct, Training Strategy=Inference-only2026.05 | 61.5 | 32.3 | |
| MEM1Backbone=Qwen3-4B-Instruct, Training Strategy=RL-trained2026.05 | 61.5 | 50.2 | |
| ReActBackbone=Qwen3-4B-Instruct, Training Strategy=Inference-only2026.05 | 60.5 | 12.6 | |
| Direct-Action (RL)Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 56 | 35.8 | |
| ReAct (RL)Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 46.5 | 34.2 | |
| MEM1Backbone=Qwen2.5-3B-Instruct, Training Strategy=RL-trained2026.05 | 29.5 | 62.9 | |
| ReActBackbone=Qwen2.5-3B-Instruct, Training Strategy=Inference-only2026.05 | 23 | 37.6 | |
| Base ModelBackbone=Qwen2.5-3B-Instruct, Training Strategy=Inference-only2026.05 | 4 | 96.1 |