Jailbreaking on AdvBench (ASR, Toxicity, WASR, W-Toxicity Scores)
61.2ASRwide-net-casting jailbreak method
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| wide-net-casting jailbreak methodTarget Model=GLM-4V-Plus, Evaluation Scenario=wide-net-casting2026.05 | 61.2 | 58 | — | — | |
| wide-net-casting jailbreak methodTarget Model=GPT-4o, Evaluation Scenario=wide-net-casting2026.05 | 55.2 | 52.3 | — | — | |
| Baseline (MLAI + PixArt-α)Target Model=GLM-4V-Plus, Evaluation Scenario=wide-net-casting2026.05 | 52.1 | 48.9 | — | — | |
| wide-net-casting jailbreak methodTarget Model=Qwen-VL-Max, Evaluation Scenario=wide-net-casting2026.05 | 51.2 | 48.1 | — | — | |
| Baseline (MLAI + PixArt-α)Target Model=GPT-4o, Evaluation Scenario=wide-net-casting2026.05 | 47.7 | 43.3 | — | — | |
| wide-net-casting jailbreak methodTarget Model=Gemini-1.5-Pro, Evaluation Scenario=wide-net-casting2026.05 | 46.1 | 42.4 | — | — | |
| Baseline (MLAI + PixArt-α)Target Model=Qwen-VL-Max, Evaluation Scenario=wide-net-casting2026.05 | 43.6 | 39.2 | — | — | |
| Baseline (MLAI + PixArt-α)Target Model=Gemini-1.5-Pro, Evaluation Scenario=wide-net-casting2026.05 | 38.9 | 34.1 | — | — | |
| Baseline (MLAI + PixArt-α)Target Model=Aggregate (All Models), Evaluation Scenario=wide-net-casting2026.05 | — | — | 69.5 | 65.3 | |
| wide-net-casting jailbreak methodTarget Model=Aggregate (All Models), Evaluation Scenario=wide-net-casting2026.05 | — | — | 86.8 | 79.9 |