Privacy-preserving Agent Interaction on Privacy Defense Evaluation Before Attack
75.9PPCDI
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CDIBackbone=gpt-4.1-mini, Instructor model=Averaged over five models (Qwen3-4B, Qwen3-4B-SafeRL, gpt-oss-20B, gpt-oss-safeguard-20B, gpt-4.1-mini)2026.03 | 75.9 | 86.9 | 82.8 | |
| PromptingBackbone=gpt-4.1-mini2026.03 | 48.1 | 73.1 | 65 | |
| GuardingBackbone=gpt-4.1-mini, Guard model=Averaged over five models (Qwen3-4B, Qwen3-4B-SafeRL, gpt-oss-20B, gpt-oss-safeguard-20B, gpt-4.1-mini)2026.03 | 47 | 82 | 70 | |
| N/ABackbone=gpt-4.1-mini2026.03 | 35.5 | 81.2 | 66.1 |