Safety Judging on Human-labeled safety evaluation set
94.23HarmBench Accuracy12B Curriculum Judge
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| 12B Curriculum JudgeGroup=Ours (Curriculum)2026.06 | 94.23 | 94.12 | 94.88 | 0.76 | |
| 12B Judge (Fixed)Group=Ours (Fixed)2026.06 | 93.16 | 91.72 | 92.37 | 1.44 | |
| Qwen2.5-14B-itGroup=BASE2026.06 | 93.08 | 85.19 | 92.88 | 7.89 | |
| 12B Judge (Dynamic)Group=Ours (Dynamic)2026.06 | 92.84 | 89.84 | 89.24 | 3.6 | |
| gpt-oss-safeguard-20BGroup=REASONING2026.06 | 92.69 | 92.23 | 94.62 | 2.39 | |
| gemma-3-12b-itGroup=BASE2026.06 | 91.35 | 85.19 | 85.96 | 6.16 | |
| gpt-oss-20BGroup=REASONING2026.06 | 90.96 | 92.23 | 92.31 | 1.35 | |
| Qwen3-30B-A3B-ThinkingGroup=REASONING2026.06 | 85 | 92.5 | 89.81 | 7.5 | |
| Llama-3.1-8B-itGroup=BASE2026.06 | 77.5 | 67.31 | 81.92 | 14.61 | |
| Llama-Guard-3-8BGroup=GUARD2026.06 | 75 | 59.62 | 84.23 | 24.61 | |
| GuardReasoner-3BGroup=GUARD2026.06 | 65.19 | 87.31 | 63.65 | 23.66 |