Agent behavioral safety on InjecAgent
95.1Safety RateThought-Aligner-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Thought-Aligner-7BCore LLM=DeepSeek-V32025.05 | 95.1 | 79.7 | |
| Thought-Aligner-7BCore LLM=Llama-3.3-70B2025.05 | 95 | 78.9 | |
| Thought-Aligner-1.5BCore LLM=DeepSeek-V32025.05 | 94.6 | 86.7 | |
| GuardAgentCore LLM=DeepSeek-V32025.05 | 94.3 | 85.5 | |
| Thought-Aligner-1.5BCore LLM=Llama-3.3-70B2025.05 | 94.3 | 85.1 | |
| Self-ReflectionCore LLM=Llama-3.3-70B2025.05 | 91.9 | 89.6 | |
| ShieldAgentCore LLM=DeepSeek-V32025.05 | 87.3 | 86.6 | |
| AthenaCore LLM=DeepSeek-V32025.05 | 87 | 72.8 | |
| GuardAgentCore LLM=Llama-3.3-70B2025.05 | 83.9 | 63.3 | |
| Self-ReflectionCore LLM=DeepSeek-V32025.05 | 83.5 | 86.9 | |
| No-GuardRailCore LLM=DeepSeek-V32025.05 | 69.9 | 86.4 | |
| ShieldAgentCore LLM=Llama-3.3-70B2025.05 | 63.8 | 88.9 | |
| AthenaCore LLM=Llama-3.3-70B2025.05 | 59.2 | 74.2 | |
| No-GuardRailCore LLM=Llama-3.3-70B2025.05 | 32.1 | 85.4 |