Agent Behavioral Safety on AgentHarm
90.6Safety RateThought-Aligner-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Thought-Aligner-7BCore LLM=Llama-3.3-70B2025.05 | 90.6 | 30 | |
| GuardAgentCore LLM=Llama-3.3-70B2025.05 | 90.4 | 39.8 | |
| Thought-Aligner-7BCore LLM=DeepSeek-V32025.05 | 89.3 | 36.5 | |
| Thought-Aligner-1.5BCore LLM=Llama-3.3-70B2025.05 | 88.8 | 34 | |
| Thought-Aligner-1.5BCore LLM=DeepSeek-V32025.05 | 88.7 | 33.2 | |
| AthenaCore LLM=Llama-3.3-70B2025.05 | 88 | 50.6 | |
| GuardAgentCore LLM=DeepSeek-V32025.05 | 87 | 46 | |
| Self-ReflectionCore LLM=Llama-3.3-70B2025.05 | 86.6 | 68.1 | |
| AthenaCore LLM=DeepSeek-V32025.05 | 81.3 | 51.2 | |
| Self-ReflectionCore LLM=DeepSeek-V32025.05 | 80.9 | 53.4 | |
| ShieldAgentCore LLM=Llama-3.3-70B2025.05 | 64.2 | 41.9 | |
| ShieldAgentCore LLM=DeepSeek-V32025.05 | 63.4 | 54.8 | |
| No-GuardRailCore LLM=Llama-3.3-70B2025.05 | 61.8 | 84 | |
| No-GuardRailCore LLM=DeepSeek-V32025.05 | 42.8 | 85 |