Online Learning of Adversarial Robustness on IH-Challenge (composite category)
75.2Final ScoreSDPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SDPOframework=CLaaS, Amax=50, backbone=Qwen3-8B2026.06 | 75.2 | 61.2 | 4.2 | |
| PPOframework=CLaaS, Amax=25, backbone=Qwen3-8B2026.06 | 49 | 37.6 | 5.4 | |
| REINFORCE++framework=CLaaS, Amax=25, backbone=Qwen3-8B2026.06 | 43.9 | 37 | 8.3 | |
| Baselineadaptation=unadapted, backbone=Qwen3-8B2026.06 | 27.2 | — | — | |
| ICLmode=in-context learning, backbone=Qwen3-8B2026.06 | 24.1 | 28.3 | 8.9 |