Guarded Agent Evaluation on ASB latest (DPI)
95.25ASRReAct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ReActAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=None2026.01 | 95.25 | 18.75 | |
| ReAct-sandwich defenseAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=Sandwich2026.01 | 91.25 | 25.25 | |
| ReActAgent Backbone=GPT-4o, Guardrail Model=None2026.01 | 82.25 | 12.5 | |
| ReAct-sandwich defenseAgent Backbone=GPT-4o, Guardrail Model=Sandwich2026.01 | 66 | 28.05 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=Safiron2026.01 | 62.25 | 8.25 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=Safiron2026.01 | 62.25 | 10.5 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=ShieldAgent-THU2026.01 | 44.5 | 9.25 | |
| ReAct-llamafirewallAgent Backbone=GPT-4o, Guardrail Model=GPT-4o-mini2026.01 | 33.28 | 10.75 | |
| ReAct-llamafirewallAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=GPT-4o-mini2026.01 | 24.22 | 14.12 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=ShieldAgent-THU2026.01 | 18.75 | 9.5 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=TS-Guard2026.01 | 7.25 | 30 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=TS-Guard2026.01 | 6.76 | 18.87 | |
| ReAct-llamafirewallAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=TS-Guard2026.01 | 5.75 | 4 | |
| ReAct-llamafirewallAgent Backbone=GPT-4o, Guardrail Model=TS-Guard2026.01 | 5.5 | 2.75 |