Scientific Safety Evaluation on SciSafetyBench
4.93Safety ScoreSciTrace
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SciTraceModel=GPT-4o2026.06 | 4.93 | 95 | 94.7 | 2.22 | 2.75 | 3.82 | |
| SciTraceModel=DeepSeek-V32026.06 | 4.91 | 94 | 93.8 | 2.18 | 2.68 | 3.75 | |
| SciTraceModel=Qwen2.5-72B2026.06 | 4.89 | 93 | 92.5 | 2.15 | 2.65 | 3.72 | |
| SciTraceModel=Llama-3.1-70B2026.06 | 4.87 | 92 | 91.2 | 2.12 | 2.62 | 3.68 | |
| SafeScientistModel=GPT-4o2026.06 | 4.83 | 90 | 81.2 | 2.1 | 2.62 | 3.62 | |
| SafeScientistModel=DeepSeek-V32026.06 | 4.78 | 88 | 79.5 | 2.05 | 2.52 | 3.52 | |
| SafeScientistModel=Qwen2.5-72B2026.06 | 4.75 | 87 | 78.1 | 2.02 | 2.5 | 3.5 | |
| SafeScientistModel=Llama-3.1-70B2026.06 | 4.72 | 85 | 76.3 | 2 | 2.48 | 3.47 | |
| Bare LLMModel=GPT-4o2026.06 | 2.5 | 0 | 45.5 | 1.92 | 2.02 | 3.3 | |
| Bare LLMModel=DeepSeek-V32026.06 | 2.4 | 2 | 42 | 1.82 | 1.87 | 3.18 | |
| Bare LLMModel=Qwen2.5-72B2026.06 | 2.38 | 0 | 40.2 | 1.8 | 1.85 | 3.15 | |
| Bare LLMModel=Llama-3.1-70B2026.06 | 2.35 | 0 | 38.5 | 1.78 | 1.82 | 3.1 |