Safety Control on DialoGPT large
0.647Safety-Quality ScoreSafeCtrl-RL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SafeCtrl-RLModel Type=Our Proposed2026.05 | 0.647 | 0.36 | |
| minimalModel Type=Hand-crafted2026.05 | 0.589 | 0.302 | |
| performance_tieredModel Type=Hand-crafted2026.05 | 0.578 | 0.29 | |
| raw_historyModel Type=Hand-crafted2026.05 | 0.567 | 0.28 | |
| contrast_learningModel Type=Hand-crafted2026.05 | 0.528 | 0.241 | |
| hybridModel Type=Hand-crafted2026.05 | 0.526 | 0.239 | |
| ai_onlyModel Type=Hand-crafted2026.05 | 0.501 | 0.213 | |
| ai_enhancedModel Type=Hand-crafted2026.05 | 0.497 | 0.21 | |
| best_worst_recentModel Type=Hand-crafted2026.05 | 0.493 | 0.205 | |
| progressiveModel Type=Hand-crafted2026.05 | 0.477 | 0.19 | |
| smart_adaptiveModel Type=Hand-crafted2026.05 | 0.439 | 0.152 | |
| trajectory_learningModel Type=Hand-crafted2026.05 | 0.408 | 0.121 | |
| OPROModel Type=Prompt Optimisation2026.05 | 0.083 | -0.204 | |
| DynamicRetrievalModel Type=Prompt Optimisation2026.05 | 0.082 | -0.205 | |
| GRIPSModel Type=Prompt Optimisation2026.05 | 0.08 | -0.207 | |
| TextGradientModel Type=Prompt Optimisation2026.05 | 0.077 | -0.21 | |
| EvolutionaryModel Type=Prompt Optimisation2026.05 | 0.027 | -0.26 |