Ethical Robustness Evaluation on AMST (post-stress)
0.93Mean RobustnessGPT-4o
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GPT-4oModel=GPT-4o2026.04 | 0.93 | 0.02 | 0.25 | 36.8 | 0.91 | 1 | |
| DeepSeek-v3Model=DeepSeek-v32026.04 | 0.68 | 0.03 | 0.14 | 25.9 | 0.65 | 2 | |
| LLaMA-3-8BModel=LLaMA-3-8B2026.04 | 0.54 | 0.03 | — | — | 0.51 | 3 |