Continuous improvement of language agents on DS-1000 StreamBench (32 held-out episodes)
56.3Pass@4Memory w/ MEMOPILOT
Evaluation Results
| Method | Links | |
|---|---|---|
| Memory w/ MEMOPILOTMemory Strategy=MEMOPILOT, Execution Agent=Qwen2.5-14B-Instruct2026.06 | 56.3 | |
| Full HistoryMemory Strategy=Full History, Execution Agent=Qwen2.5-14B-Instruct2026.06 | 52.5 | |
| No MemoryMemory Strategy=No Memory, Execution Agent=Qwen2.5-14B-Instruct2026.06 | 50 | |
| Memory w/ DeepSeek-V3.2Memory Strategy=DeepSeek-V3.2, Execution Agent=Qwen2.5-14B-Instruct2026.06 | 50 | |
| Memory w/ Qwen2.5-14BMemory Strategy=Qwen2.5-14B, Execution Agent=Qwen2.5-14B-Instruct2026.06 | 48.8 |