Safety Veto Detection on ProMedical-Bench Safety dimension (S3)
91.5PrecisionProMedical-RM (Qwen3)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ProMedical-RM (Qwen3)Model Category=Ours, Medical-specific=true2026.04 | 91.5 | 86.8 | 89.09 | |
| ProMedical-RM (Llama)Model Category=Ours, Medical-specific=true2026.04 | 89.4 | 85.1 | 87.2 | |
| DeepSeek-R1Model Category=Open-Source Generative, Medical-specific=false2026.04 | 81.5 | 76.28 | 78.8 | |
| Qwen3-235B-ThinkingModel Category=Open-Source Generative, Medical-specific=false2026.04 | 80.15 | 76.1 | 78.07 | |
| GPT-5Model Category=Closed-Source Generative, Medical-specific=false2026.04 | 79.24 | 73.85 | 76.45 | |
| Gemini-3-ProModel Category=Closed-Source Generative, Medical-specific=false2026.04 | 68.5 | 60.25 | 64.11 | |
| Qwen3-8BModel Category=Open-Source Generative, Medical-specific=false2026.04 | 66.4 | 63.8 | 65.07 | |
| PairRM-LLaMA3-8BModel Category=Reward Models, Medical-specific=false2026.04 | 62.45 | 59.8 | 61.1 | |
| HuatuoGPT-o1Model Category=Open-Source Generative, Medical-specific=true2026.04 | 61.2 | 55.5 | 58.21 | |
| medical_o1_verifierModel Category=Reward Models, Medical-specific=true2026.04 | 55.3 | 50.8 | 52.95 |