Guarded Agent Evaluation on AgentHarm latest (full)
97.16Refusal RateReAct-llamafirewall
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ReAct-llamafirewallAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=TS-Guard2026.01 | 97.16 | 6.6 | |
| ReAct-llamafirewallAgent Backbone=GPT-4o, Guardrail Model=TS-Guard2026.01 | 96.59 | 4.28 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=TS-Guard2026.01 | 95.45 | 6.83 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=ShieldAgent-THU2026.01 | 94.88 | 6.31 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=TS-Guard2026.01 | 94.32 | 6.03 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=ShieldAgent-THU2026.01 | 92.1 | 7.12 | |
| ReAct-llamafirewallAgent Backbone=GPT-4o, Guardrail Model=GPT-4o-mini2026.01 | 80.87 | 12.64 | |
| ReAct-sandwich defenseAgent Backbone=GPT-4o, Guardrail Model=Sandwich2026.01 | 77.84 | 13.74 | |
| ReAct-TS-FlowAgent Backbone=GPT-4o, Guardrail Model=Safiron2026.01 | 73.86 | 18.91 | |
| ReAct-llamafirewallAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=GPT-4o-mini2026.01 | 70.45 | 22.99 | |
| ReAct-TS-FlowAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=Safiron2026.01 | 67.04 | 19.56 | |
| ReActAgent Backbone=GPT-4o, Guardrail Model=None2026.01 | 62.5 | 23.53 | |
| ReAct-sandwich defenseAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=Sandwich2026.01 | 50 | 33.4 | |
| ReActAgent Backbone=Qwen2.5-14B-Instruct, Guardrail Model=None2026.01 | 42.04 | 34.14 |