Safety Evaluation on HONEST
1.8ScoreUniversal Self-consistency
Evaluation Results
| Method | Links | |
|---|---|---|
| Universal Self-consistencyGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 1.8 | |
| Self-refineGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 1.3 | |
| Zero-shot-CoTGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 1.1 | |
| MADGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 0.9 | |
| ExpertPromptingGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 0.8 | |
| MEPGroup=Baselines, Base Model=GPT-3.5-Turbo2025.07 | 0.7 | |
| MetaCrit w/o ϕ↑Group=Ablation, Base Model=GPT-3.5-Turbo2025.07 | 0 | |
| MetaCrit w/o ϕ↓Group=Ablation, Base Model=GPT-3.5-Turbo2025.07 | 0 | |
| MetaCrit + GPT-3.5-TurboGroup=Ours, Base Model=GPT-3.5-Turbo2025.07 | 0 | |
| MetaCrit + DeepSeek-v3Group=Ours, Base Model=DeepSeek-v32025.07 | 0 | |
| MetaCrit + Claude-3.5-SonnetGroup=Ours, Base Model=Claude-3.5-Sonnet2025.07 | 0 | |
| MetaCrit + GPT-4oGroup=Ours, Base Model=GPT-4o2025.07 | 0 |