Guarded Agent Evaluation on AgentDojo full latest
0.5616ASRReAct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ReActAgent Backbone=GPT-4o, Guardrail Model=None2026.01 | 0.5616 | 0.2687 | |
| ReAct-sandwich defenseAgent Backbone=GPT-4o, Guardrail Model=Sandwich2026.01 | 0.5412 | 0.2895 | |
| ReAct-sandwich defenseAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=Sandwich2026.01 | 0.1854 | 0.4269 | |
| ReActAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=None2026.01 | 0.1759 | 0.4257 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=Safiron2026.01 | 0.0768 | 0.2239 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=Safiron2026.01 | 0.0685 | 0.2582 | |
| ReAct-llamafirewallAgent Backbone=GPT-4o, Guardrail Model=GPT-4o-mini2026.01 | 0.0302 | 0.2447 | |
| ReAct-llamafirewallAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=GPT-4o-mini2026.01 | 0.0284 | 0.3382 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=TS-Guard2026.01 | 0.0179 | 0.4272 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=ShieldAgent-THU2026.01 | 0.0135 | 0.2486 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=TS-Guard2026.01 | 0.0116 | 0.4278 | |
| ReAct-llamafirewallAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=TS-Guard2026.01 | 0.0105 | 0.3646 | |
| ReAct-llamafirewallAgent Backbone=GPT-4o, Guardrail Model=TS-Guard2026.01 | 0.0095 | 0.2079 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=ShieldAgent-THU2026.01 | 0.0093 | 0.2644 |