LLM Unlearning on MMLU
100AccuracyLLaMA-3.1-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLaMA-3.1-8BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 100 | — | — | |
| Yi-34B + Activation-based DetectionModel=Yi-34B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 99.86 | — | — | |
| Zephyr-7BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 99.86 | — | — | |
| Yi-34BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 99.86 | — | — | |
| Zephyr-7BUnlearning method=NPO, Detection source=Text-based2025.06 | 99.86 | — | — | |
| LLaMA-3.1-8B + Activation-based DetectionModel=LLaMA-3.1-8B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 99.72 | — | — | |
| LLaMA-3.1-8BUnlearning method=NPO, Detection source=Text-based2025.06 | 99.72 | — | — | |
| Qwen2.5-14BUnlearning method=NPO, Detection source=Text-based2025.06 | 99.72 | — | — | |
| Qwen2.5-14BUnlearning Method=NPO, Input Features=RMS-normalized final-layer activations, Unlearning Target Dataset=WMDP2025.06 | 99.44 | — | — | |
| Yi-34BUnlearning method=NPO, Detection source=Text-based2025.06 | 98.87 | — | — | |
| Zephyr-7B + Activation-based DetectionModel=Zephyr-7B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 98.59 | — | — | |
| Qwen2.5-14B + Activation-based DetectionModel=Qwen2.5-14B, Unlearning method=RMU, Feature type=RMS-normalized final-layer activations, Classifier=two-layer MLP, Unlearning target=WMDP2025.06 | 98.31 | — | — | |
| Base ModelBackbone=Llama-3-8B2025.02 | 62.19 | — | — | |
| RMUBackbone=Llama-3-8B2025.02 | 61.13 | — | — | |
| NPOGDRBackbone=Llama-3-8B2025.02 | 59.96 | — | — | |
| GAGDRBackbone=Llama-3-8B2025.02 | 59.84 | — | — | |
| Base ModelBackbone=Mistral-7B2025.02 | 59.13 | — | — | |
| NPOKLRBackbone=Llama-3-8B2025.02 | 56.15 | — | — | |
| GAKLRBackbone=Llama-3-8B2025.02 | 55.7 | — | — | |
| RMUBackbone=Mistral-7B2025.02 | 49.91 | — | — | |
| NPOKLRBackbone=Mistral-7B2025.02 | 49.16 | — | — | |
| GAKLRBackbone=Mistral-7B2025.02 | 47.02 | — | — | |
| NPOGDRBackbone=Mistral-7B2025.02 | 42.81 | — | — | |
| TVBackbone=Llama-3-8B2025.02 | 34.2 | — | — | |
| GABackbone=Llama-3-8B2025.02 | 28.56 | — | — | |
| TVBackbone=Mistral-7B2025.02 | 25.8 | — | — | |
| NPOBackbone=Mistral-7B2025.02 | 25.51 | — | — | |
| GABackbone=Mistral-7B2025.02 | 24.65 | — | — | |
| NPOBackbone=Llama-3-8B2025.02 | 22.95 | — | — | |
| GAGDRBackbone=Mistral-7B2025.02 | 22.74 | — | — | |
| GradAscentretain free=true, Time=173s2026.04 | — | 25.5 | 23 | |
| GradDiffretain free=false, Time=280s2026.04 | — | 58.5 | 59.3 | |
| MC-WIN-Uretain free=true, Time=50679s2026.04 | — | 55.1 | 56.5 | |
| NPOretain free=false, Time=1178s2026.04 | — | 59.1 | 59.2 | |
| Original model2026.04 | — | 59.2 | — | |
| RMUretain free=false, Time=1420s2026.04 | — | 59.2 | 59.3 | |
| SimNPOretain free=false, Time=545s2026.04 | — | 59.5 | 59.5 |