Unlearning Detection on UltraChat
99.86AccuracyYi-34B + Activation-based Detection
Evaluation Results
| Method | Links | |
|---|---|---|
| Yi-34B + Activation-based DetectionModel=Yi-34B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 99.86 | |
| Zephyr-7BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 99.86 | |
| Yi-34BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 99.86 | |
| LLaMA-3.1-8BUnlearning method=NPO, Detection source=Text-based2025.06 | 99.72 | |
| LLaMA-3.1-8B + Activation-based DetectionModel=LLaMA-3.1-8B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 99.44 | |
| LLaMA-3.1-8BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 99.44 | |
| Qwen2.5-14BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 99.44 | |
| Qwen2.5-14BUnlearning method=NPO, Detection source=Text-based2025.06 | 99.44 | |
| Zephyr-7BUnlearning method=NPO, Detection source=Text-based2025.06 | 99.16 | |
| Zephyr-7B + Activation-based DetectionModel=Zephyr-7B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 99.15 | |
| Qwen2.5-14B + Activation-based DetectionModel=Qwen2.5-14B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 99.15 | |
| Yi-34BUnlearning method=NPO, Detection source=Text-based2025.06 | 99.15 |