Privacy and Completeness evaluation on CIMemories GPT-5.5 low reasoning labels (held out)
2.6Violation RateGPT-5.5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5.5Reasoning mode=high2026.06 | 2.6 | 46.3 | |
| GPT-5.4-miniReasoning mode=high2026.06 | 2.9 | 36.9 | |
| Gemini-3.1-Flash-LiteReasoning mode=high2026.06 | 9.1 | 44 | |
| Gemini-3.1-ProReasoning mode=high2026.06 | 16.7 | 60.1 | |
| Nemotron-3-Nano-4B + RL (PrivacyAlign)Reasoning mode=thinking, RL training=annotation-conditioned reward2026.06 | 25.4 | 35.6 | |
| Nemotron-3-Nano-4BReasoning mode=thinking2026.06 | 33.4 | 32.6 | |
| Qwen3-8B + RL (PrivacyAlign)Reasoning mode=thinking, RL training=annotation-conditioned reward2026.06 | 38.3 | 49 | |
| Qwen3-8BReasoning mode=thinking2026.06 | 44.7 | 47.4 | |
| Qwen3-4B + RL (PrivacyAlign)Reasoning mode=thinking, RL training=annotation-conditioned reward2026.06 | 49.6 | 51.4 | |
| Qwen3-4BReasoning mode=thinking2026.06 | 51 | 44.4 |