Agent Behavioral Safety on AgentDojo
97.1Safety RateThought-Aligner-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Thought-Aligner-7BCore LLM=DeepSeek-V32025.05 | 97.1 | 34.4 | |
| Thought-Aligner-1.5BCore LLM=DeepSeek-V32025.05 | 96.8 | 38.5 | |
| AthenaCore LLM=DeepSeek-V32025.05 | 94.9 | 56.3 | |
| Thought-Aligner-7BCore LLM=Llama-3.3-70B2025.05 | 93 | 31.6 | |
| Thought-Aligner-1.5BCore LLM=Llama-3.3-70B2025.05 | 92.9 | 45.7 | |
| Self-ReflectionCore LLM=Llama-3.3-70B2025.05 | 92.7 | 46.9 | |
| AthenaCore LLM=Llama-3.3-70B2025.05 | 92 | 45.8 | |
| Self-ReflectionCore LLM=DeepSeek-V32025.05 | 90.3 | 51 | |
| GuardAgentCore LLM=DeepSeek-V32025.05 | 89.5 | 44.8 | |
| GuardAgentCore LLM=Llama-3.3-70B2025.05 | 88.5 | 47.9 | |
| ShieldAgentCore LLM=Llama-3.3-70B2025.05 | 83.1 | 61.5 | |
| ShieldAgentCore LLM=DeepSeek-V32025.05 | 74.3 | 62.5 | |
| No-GuardRailCore LLM=DeepSeek-V32025.05 | 67 | 64.6 | |
| No-GuardRailCore LLM=Llama-3.3-70B2025.05 | 53.4 | 77.7 |