Language Model Unlearning on TOFU (Forget10)
100Forget Quality (FQ)Retain90
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Retain90Backbone=Phi-1.5B2025.09 | 100 | 52 | 43 | 91 | — | |
| GD+SineBackbone=Phi-1.5B, Method Category=Parameter-Efficient Methods, Params (%)=1.62025.09 | 94.3 | 52 | 22 | 90 | — | |
| ME+GD (LoRA)Backbone=Phi-1.5B, Method Category=Parameter-Efficient Methods, Params (%)=1.62025.09 | 78.6 | 52 | 14 | 93 | — | |
| GD+TanhBackbone=Phi-1.5B, Method Category=Parameter-Efficient Methods, Params (%)=1.62025.09 | 34.2 | 49 | 28 | 85 | — | |
| before unlearningBackbone Model=Qwen-3-8B2026.05 | 0.9025 | 76.03 | — | — | 88.47 | |
| before unlearningBackbone Model=Llama-2-7B2026.05 | 0.8122 | 62.19 | — | — | 83.64 | |
| before unlearningBackbone Model=Phi-1.52026.05 | 0.5796 | 52.88 | — | — | 43.17 | |
| NPOBackbone Model=Llama-2-7B, Unlearning Granularity=Sequence-Level2026.05 | 0.2139 | 60.05 | — | — | 47.12 | |
| GABackbone Model=Qwen-3-8B, Unlearning Granularity=Sequence-Level2026.05 | 0.1739 | 57.3 | — | — | 45.42 | |
| RMUBackbone Model=Llama-2-7B, Unlearning Granularity=Sequence-Level2026.05 | 0.1715 | 51.74 | — | — | 29.65 | |
| NPOBackbone Model=Qwen-3-8B, Unlearning Granularity=Sequence-Level2026.05 | 0.1684 | 55.97 | — | — | 49.3 | |
| NPOBackbone Model=Phi-1.5, Unlearning Granularity=Sequence-Level2026.05 | 0.1636 | 47.92 | — | — | 24.2 | |
| GABackbone Model=Llama-2-7B, Unlearning Granularity=Sequence-Level2026.05 | 0.1632 | 50.74 | — | — | 41.79 | |
| GABackbone Model=Phi-1.5, Unlearning Granularity=Sequence-Level2026.05 | 0.1425 | 32.64 | — | — | 19.63 | |
| RMUBackbone Model=Phi-1.5, Unlearning Granularity=Sequence-Level2026.05 | 0.1404 | 49.27 | — | — | 23.75 | |
| S-GABackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.1368 | 34.58 | — | — | 28.05 | |
| RMUBackbone Model=Qwen-3-8B, Unlearning Granularity=Sequence-Level2026.05 | 0.1365 | 58.82 | — | — | 40.35 | |
| WGABackbone Model=Phi-1.5, Unlearning Granularity=Sequence-Level2026.05 | 0.1328 | 50.86 | — | — | 34.18 | |
| S-NPOBackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.1323 | 56.37 | — | — | 61.39 | |
| S-RMUBackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.1296 | 60.42 | — | — | 54.54 | |
| T-RMUBackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.1273 | 51.06 | — | — | 27.19 | |
| S-RMUBackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.1258 | 50.97 | — | — | 24.39 | |
| T-GABackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.1255 | 35.29 | — | — | 28.47 | |
| WGABackbone Model=Llama-2-7B, Unlearning Granularity=Sequence-Level2026.05 | 0.1255 | 64.39 | — | — | 64.18 | |
| S-WGABackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.1243 | 52.37 | — | — | 35.12 | |
| S-GABackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.1216 | 60.65 | — | — | 65.03 | |
| T-GABackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.1209 | 63.42 | — | — | 66.23 | |
| S-WGABackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.115 | 60.12 | — | — | 65.74 | |
| T-NPOBackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.1138 | 58.12 | — | — | 63.89 | |
| WGABackbone Model=Qwen-3-8B, Unlearning Granularity=Sequence-Level2026.05 | 0.1072 | 58.24 | — | — | 57.61 | |
| S-NPOBackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.1023 | 48.21 | — | — | 27.65 | |
| T-NPOBackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.1004 | 49.36 | — | — | 29.85 | |
| S-WGABackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.0955 | 65.12 | — | — | 65.24 | |
| T-WGABackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.0949 | 66.5 | — | — | 67.91 | |
| T-WGABackbone Model=Phi-1.5, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.0923 | 52.42 | — | — | 36.14 | |
| S-GABackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.0917 | 53.21 | — | — | 57.22 | |
| S-RMUBackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.0835 | 59.14 | — | — | 41.96 | |
| T-NPOBackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.0817 | 60.89 | — | — | 65.83 | |
| T-RMUBackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.0814 | 59.36 | — | — | 58.39 | |
| T-WGABackbone Model=Qwen-3-8B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.0723 | 60.53 | — | — | 68.57 | |
| T-GABackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.0715 | 57.13 | — | — | 63 | |
| T-RMUBackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Hard Selection2026.05 | 0.0709 | 58.62 | — | — | 43.38 | |
| S-NPOBackbone Model=Llama-2-7B, Unlearning Granularity=Token-Level, Selection Strategy=Soft Weighting2026.05 | 0.0664 | 55.39 | — | — | 59.75 | |
| NPOBackbone=Phi-1.5B, Method Category=Full Fine-tuning Methods, Params (%)=100.02025.09 | 0.0026 | 37 | 45 | 45 | — | |
| GD+FILABackbone=Phi-1.5B, Method Category=Parameter-Efficient Methods, Params (%)=1.62025.09 | 0.0002 | 0 | 12 | 11 | — | |
| GDBackbone=Phi-1.5B, Method Category=Full Fine-tuning Methods, Params (%)=100.02025.09 | 0 | 36 | 37 | 41 | — | |
| LoKUBackbone=Phi-1.5B, Method Category=Parameter-Efficient Methods, Params (%)=1.62025.09 | 0 | 51 | 26 | 75 | — | |
| GABackbone=Phi-1.5B, Method Category=Full Fine-tuning Methods, Params (%)=100.02025.09 | 0 | 0 | 1 | 1 | — | |
| KLBackbone=Phi-1.5B, Method Category=Full Fine-tuning Methods, Params (%)=100.02025.09 | 0 | 0 | 1 | 1 | — | |
| GD+LoRABackbone=Phi-1.5B, Method Category=Parameter-Efficient Methods, Params (%)=1.62025.09 | 0 | 28 | 85 | 45 | — | |
| DPOBackbone=Phi-1.5B, Method Category=Full Fine-tuning Methods, Params (%)=100.02025.09 | 0 | 48 | 41 | 67 | — | |
| GA+FILABackbone=Phi-1.5B, Method Category=Parameter-Efficient Methods, Params (%)=1.62025.09 | 0 | 0 | 0 | 0 | — | |
| IHLBackbone=Phi-1.5B, Method Category=Full Fine-tuning Methods, Params (%)=100.02025.09 | 0 | 51 | 53 | 76 | — | |
| OriginalBackbone=Phi-1.5B2025.09 | 0 | 52 | 93 | 92 | — |