Agent Safety on ASSEBench
92.04AccuracyDRAFT (Qwen3Guard-Gen-4B)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DRAFT (Qwen3Guard-Gen-4B)Backbone=Qwen3Guard-Gen-4B, Adaptation=DRAFT2026.02 | 92.04 | 90.55 | — | 88.77 | |
| DRAFT (Qwen3-8B)Backbone=Qwen3-8B, Adaptation=DRAFT2026.02 | 91.57 | 89.75 | — | 87.19 | |
| Gemini-2Evaluation Setup=+AA∆(%)2026.05 | 91.5 | 91.44 | — | — | |
| DRAFT (Qwen3-4B-Instruct-2507)Backbone=Qwen3-4B-Instruct-2507, Adaptation=DRAFT2026.02 | 91.38 | 88.98 | — | 86.04 | |
| ExtractorBackbone=Llama-3.1-8B-Instruct2026.02 | 90.81 | 88.09 | 96.06 | 81.33 | |
| DRAFTBackbone=Qwen2.5-7B-Instruct2026.02 | 89.97 | 87.23 | 93.18 | 82 | |
| DRAFT (Llama-3.1-8B)Backbone=Llama-3.1-8B, Adaptation=DRAFT2026.02 | 89.72 | 86.91 | — | 84.62 | |
| DRAFTBackbone=Llama-3.1-8B-Instruct2026.02 | 89.69 | 86.25 | 97.48 | 77.33 | |
| QwQ-32BEvaluation Setup=+AA∆(%)2026.05 | 89.63 | 90.09 | — | — | |
| GPT-4.1Evaluation Setup=+AA∆(%)2026.05 | 89.12 | 88.37 | — | — | |
| Claude-3.5Evaluation Setup=+AA∆(%)2026.05 | 89.02 | 89.44 | — | — | |
| ExtractorBackbone=Qwen2.5-7B-Instruct2026.02 | 88.86 | 86.01 | 90.44 | 82 | |
| Deepseek v3Evaluation Setup=+AA∆(%)2026.05 | 88.66 | 87.81 | — | — | |
| GPT-o3-miniEvaluation Setup=+AA∆(%)2026.05 | 87.99 | 86.95 | — | — | |
| DRAFTBackbone=Qwen3-4B2026.02 | 87.19 | 83.45 | 90.62 | 77.33 | |
| GPT-4oEvaluation Setup=+AA∆(%)2026.05 | 85.63 | 84.73 | — | — | |
| Qwen-2.5-32BEvaluation Setup=+AA∆(%)2026.05 | 85.19 | 85.7 | — | — | |
| DRAFTBackbone=Qwen2.5-3B-Instruct2026.02 | 84.4 | 80.69 | 83.57 | 78.01 | |
| Qwen3-4B-Instruct-2507 (SFT)Backbone=Qwen3-4B-Instruct-2507, Adaptation=SFT2026.02 | 84.22 | 77.06 | — | 68.9 | |
| SFTBackbone=Llama-3.1-8B-Instruct2026.02 | 84.12 | 77.99 | 92.66 | 67.33 | |
| ExtractorBackbone=Qwen3-4B2026.02 | 83.01 | 77.15 | 88.03 | 68.67 | |
| ShieldAgentEvaluation Setup=Origin2026.05 | 82.33 | 82.92 | — | — | |
| Qwen-2.5-7BEvaluation Setup=+AA∆(%)2026.05 | 81.53 | 80.53 | — | — | |
| Qwen3Guard-Gen-4B (SFT)Backbone=Qwen3Guard-Gen-4B, Adaptation=SFT2026.02 | 81.09 | 74.15 | — | 68.34 | |
| ExtractorBackbone=Qwen2.5-3B-Instruct2026.02 | 81.06 | 71.43 | 96.59 | 56.67 | |
| Qwen3-8B (AgentAuditor)Backbone=Qwen3-8B, Adaptation=AA2026.02 | 80.82 | 81.84 | — | 79.33 | |
| Qwen3-8B (SFT)Backbone=Qwen3-8B, Adaptation=SFT2026.02 | 80.17 | 72.27 | — | 64.45 | |
| GPT-4.1Evaluation Setup=Origin2026.05 | 79.69 | 78.17 | — | — | |
| SFTBackbone=Qwen3-4B2026.02 | 79.39 | 70.16 | 88.78 | 58 | |
| GPT-o3-miniEvaluation Setup=Origin2026.05 | 79.37 | 76.63 | — | — | |
| Claude-3.5Evaluation Setup=Origin2026.05 | 79.31 | 81.08 | — | — | |
| ExtractorBackbone=Qwen2.5-1.5B-Instruct2026.02 | 78.83 | 69.6 | 87.06 | 58.03 | |
| Llama-3.1-8B (AgentAuditor)Backbone=Llama-3.1-8B, Adaptation=AA2026.02 | 77.99 | 79.41 | — | 70.67 | |
| Deepseek v3Evaluation Setup=Origin2026.05 | 77.58 | 74.6 | — | — | |
| DRAFTBackbone=Qwen2.5-1.5B-Instruct2026.02 | 77.16 | 67.2 | 84.04 | 56.02 | |
| Llama-3.1-8B (SFT)Backbone=Llama-3.1-8B, Adaptation=SFT2026.02 | 76.56 | 63.41 | — | 57.17 | |
| QwQ-32BEvaluation Setup=Origin2026.05 | 76.3 | 78.44 | — | — | |
| ChatGPT 5.2 (API)Backbone=ChatGPT 5.2, Adaptation=API2026.02 | 74.67 | 71.6 | — | 62.56 | |
| Qwen3Guard-Gen-4B (AgentAuditor)Backbone=Qwen3Guard-Gen-4B, Adaptation=AA2026.02 | 74.19 | 72.44 | — | 67.01 | |
| SFTBackbone=Qwen2.5-1.5B-Instruct2026.02 | 73.54 | 59.23 | 83.13 | 46 | |
| Gemini-2Evaluation Setup=Origin2026.05 | 72.74 | 65.6 | — | — | |
| GPT-4oEvaluation Setup=Origin2026.05 | 72.19 | 69 | — | — | |
| SFTBackbone=Qwen2.5-7B-Instruct2026.02 | 71.59 | 52.34 | 87.5 | 37.33 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.02 | 71.31 | 52.09 | 86.15 | 37.33 | |
| Qwen3-4B-Instruct-2507 (AgentAuditor)Backbone=Qwen3-4B-Instruct-2507, Adaptation=AA2026.02 | 71.06 | 58.92 | — | 60.3 | |
| Llama-3.1-8BEvaluation Setup=+AA∆(%)2026.05 | 70.81 | 74.9 | — | — | |
| gpt-oss-120b (Vanilla)Backbone=gpt-oss-120b, Adaptation=Vanilla2026.02 | 69.52 | 67.85 | — | 60.51 | |
| Llama-Guard-3Evaluation Setup=Origin2026.05 | 68.54 | 74.62 | — | — | |
| Qwen-2.5-32BEvaluation Setup=Origin2026.05 | 65.51 | 68.37 | — | — | |
| Llama-3.1-8B (LoRA)Backbone=Llama-3.1-8B, Adaptation=LoRA2026.02 | 65.18 | 39.02 | — | 26.67 | |
| Qwen3-8B (LoRA)Backbone=Qwen3-8B, Adaptation=LoRA2026.02 | 64.76 | 57.91 | — | 57.33 | |
| Qwen3-4B-Instruct-2507 (LoRA)Backbone=Qwen3-4B-Instruct-2507, Adaptation=LoRA2026.02 | 63.79 | 65.24 | — | 81.33 | |
| Qwen3-4B-Instruct-2507 (Vanilla)Backbone=Qwen3-4B-Instruct-2507, Adaptation=Vanilla2026.02 | 63.23 | 44.57 | — | 41.09 | |
| Qwen3Guard-Gen-4B (Vanilla)Backbone=Qwen3Guard-Gen-4B, Adaptation=Vanilla2026.02 | 62.64 | 34.36 | — | 23.66 | |
| Llama-3.1-8B (Vanilla)Backbone=Llama-3.1-8B, Adaptation=Vanilla2026.02 | 61.55 | 25.69 | — | 16.08 | |
| LoRABackbone=Llama-3.1-8B-Instruct2026.02 | 60.45 | 39.83 | 54.65 | 31.33 | |
| LoRABackbone=Qwen2.5-1.5B-Instruct2026.02 | 59.61 | 18.08 | 59.26 | 10.67 | |
| Qwen3Guard-Gen-4B (LoRA)Backbone=Qwen3Guard-Gen-4B, Adaptation=LoRA2026.02 | 59.33 | 29.81 | — | 20.67 | |
| VanillaBackbone=Llama-3.1-8B-Instruct2026.02 | 58.96 | 35.98 | 50.64 | 27.91 | |
| VanillaBackbone=Qwen2.5-1.5B-Instruct2026.02 | 58.8 | 14.41 | 50.92 | 8.39 | |
| Qwen3-8B (Vanilla)Backbone=Qwen3-8B, Adaptation=Vanilla2026.02 | 58.69 | 49.87 | — | 49.85 | |
| Qwen-2.5-7BEvaluation Setup=Origin2026.05 | 57.41 | 56.16 | — | — | |
| LoRABackbone=Qwen2.5-7B-Instruct2026.02 | 56.82 | 59.95 | 48.95 | 77.33 | |
| VanillaBackbone=Qwen3-4B2026.02 | 54.2 | 57.49 | 46.63 | 74.92 | |
| VanillaBackbone=Qwen2.5-7B-Instruct2026.02 | 54.12 | 55.94 | 46.37 | 70.48 | |
| LoRABackbone=Qwen3-4B2026.02 | 51.25 | 56.36 | 45.02 | 75.33 | |
| Llama-3.1-8BEvaluation Setup=Origin2026.05 | 51.02 | 65.19 | — | — | |
| VanillaBackbone=Qwen2.5-3B-Instruct2026.02 | 50.56 | 50.44 | 43.06 | 60.87 | |
| LoRABackbone=Qwen2.5-3B-Instruct2026.02 | 50.14 | 50.42 | 43.13 | 60.67 |