Response Moderation on FlexBench
75.81Strict ScoreFlexGuard
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| FlexGuardMode=Continuous-score, Evaluation Strategy=Rubric thresholding2026.02 | 75.81 | 83.22 | 77.03 | 78.69 | 75.81 | |
| FlexGuardMode=Continuous-score, Evaluation Strategy=Calibrated thresholding2026.02 | 75.81 | 82.68 | 82.38 | 80.29 | 75.81 | |
| FlexGuardMethod=Rubric thresholding, Score Type=continuous-score2026.02 | 75.81 | 83.22 | 77.03 | 78.69 | 75.81 | |
| FlexGuardMethod=Calibrated thresholding, Score Type=continuous-score2026.02 | 75.81 | 82.68 | 82.38 | 80.29 | 75.81 | |
| BingoGuard-8BCategory=Level-thresholded moderator2026.02 | 74.8 | 78.35 | 76.61 | 76.59 | 74.8 | |
| PKU-SafeRLHF-8BCategory=Level-thresholded moderator2026.02 | 74.54 | 81.96 | 74.15 | 76.88 | 74.15 | |
| DeepSeek-R1Category=Rubric-prompted LLM2026.02 | 74.3 | 78.06 | 70.22 | 74.19 | 70.22 | |
| GPT-5Category=Rubric-prompted LLM2026.02 | 74.07 | 81.32 | 76.9 | 77.43 | 74.07 | |
| Doubao-1.8Category=Rubric-prompted LLM2026.02 | 73.53 | 81.15 | 73.72 | 76.13 | 73.53 | |
| Qwen3Guard-Gen-8BRegime=loose2026.02 | 70.59 | 82.2 | 80.69 | 77.83 | 70.59 | |
| Qwen3Guard-8B-GenCategory=Logit-thresholded moderator2026.02 | 69.16 | 81.16 | 79.52 | 76.61 | 69.16 | |
| BingoGuard-8B2026.02 | 69.05 | 79.97 | 77.18 | 75.4 | 69.05 | |
| Qwen3Guard-Gen-8BRegime=strict2026.02 | 68.97 | 79.4 | 77.46 | 75.28 | 68.97 | |
| WildGuard-7BCategory=Logit-thresholded moderator2026.02 | 66.67 | 54.55 | 74.61 | 65.28 | 54.55 | |
| LlamaGuard3-8BCategory=Logit-thresholded moderator2026.02 | 66.67 | 70.48 | 69.65 | 68.93 | 66.67 | |
| WildGuard-7B2026.02 | 66.67 | 77.54 | 74.93 | 73.04 | 66.67 | |
| PKU-SafeRLHF-8B2026.02 | 60.79 | 72.06 | 68.79 | 67.21 | 60.79 | |
| LlamaGuard3-8B2026.02 | 59.35 | 70.48 | 69.65 | 66.49 | 59.35 |