Defense Robustness on Direct Inquiry Vanilla
100Keyword Match RateSTL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| STLLLM=*Llama-2 (7B-Chat)2026.03 | 100 | 99 | 99 | |
| DPLLLM=*Llama-2 (7B-Chat)2026.03 | 100 | 100 | 99 | |
| ES2LLM=*Llama-2 (7B-Chat)2026.03 | 100 | 100 | 100 | |
| ES2LLM=Llama-3 (8B-Instruct)2026.03 | 100 | 99 | 99 | |
| ES2LLM=Qwen-2.5 (7B-Instruct)2026.03 | 100 | 100 | 98 | |
| Base ModelLLM=*Llama-2 (7B-Chat)2026.03 | 99 | 98 | 98 | |
| STLLLM=Llama-3 (8B-Instruct)2026.03 | 98 | 98 | 96 | |
| DPLLLM=Llama-3 (8B-Instruct)2026.03 | 97 | 96 | 96 | |
| DPLLLM=Qwen-2.5 (7B-Instruct)2026.03 | 97 | 94 | 94 | |
| Base ModelLLM=Llama-3 (8B-Instruct)2026.03 | 96 | 96 | 96 | |
| STLLLM=Qwen-2.5 (7B-Instruct)2026.03 | 95 | 95 | 93 | |
| Base ModelLLM=Qwen-2.5 (7B-Instruct)2026.03 | 92 | 91 | 91 | |
| ES2LLM=Mistral (7B-Instruct)2026.03 | 67 | 63 | 58 | |
| DPLLLM=Mistral (7B-Instruct)2026.03 | 47 | 37 | 36 | |
| STLLLM=Mistral (7B-Instruct)2026.03 | 42 | 38 | 34 | |
| Base ModelLLM=Mistral (7B-Instruct)2026.03 | 30 | 27 | 26 |