Malicious Agent Detection on text-based response dataset
98Alpaca ScoreBPD
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BPDStructure=Hierarchy, Base Model=GPT-4o2025.10 | 98 | 91 | 97 | 95 | |
| BPDStructure=Flat, Base Model=GPT-4o2025.10 | 96 | 92 | 94 | 94 | |
| G-SafeguardStructure=Hierarchy, Base Model=GPT-4o2025.10 | 93 | 87 | 92 | 91 | |
| G-SafeguardStructure=Flat, Base Model=GPT-4o2025.10 | 92 | 86 | 89 | 89 |