LLM Security Defense on Macro-averaged GPT-4o, Llama-3 8B, and Mistral-7B (test)
11.3Attack Success Rate (ASR)Layered Security Framework
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Layered Security FrameworkConfiguration=Proposed Framework, Layers=L1 + L2 + L32026.06 | 11.3 | 4.8 | 88.7 | 0.89 | |
| NeMo GuardrailsConfiguration=NeMo Guardrails, Mode=Default prompt-injection protections2026.06 | 35.1 | 8.4 | 64.9 | 0.88 | |
| L3 OnlyConfiguration=L3 Only, Enabled Layer=Layer 3 (Output Auditing)2026.06 | 38.6 | 6.7 | 61.4 | 0.89 | |
| L1 OnlyConfiguration=L1 Only, Enabled Layer=Layer 1 (Input Screening)2026.06 | 44.2 | 5.1 | 55.8 | 0.9 | |
| L2 OnlyConfiguration=L2 Only, Enabled Layer=Layer 2 (Privilege-Constrained Context Assembly)2026.06 | 51.8 | 1.3 | 48.2 | 0.89 | |
| UndefendedConfiguration=Undefended, Defense Layers=None2026.06 | 71.4 | 0 | 28.6 | 0.91 |