Machine Unlearning on MUSE-Books Harry Potter v1.0 (Overall)
32.13R-ForgetBase Model
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Base ModelBackbone=Llama3.2-3B2026.01 | 32.13 | 39.99 | 84.29 | 61.46 | 0 | |
| Refusal-TrainingForget set=Query-refusal response pairs (Df_QR), Context=Retention set Dr2026.01 | 31.02 | 37.75 | 75.32 | 60.48 | -6.6 | |
| NPOForget set=QA samples (Df_QA)2026.01 | 30.19 | 34.28 | 46.2 | 60.48 | -31.42 | |
| NPO + KLKL-divergence regularization=Retention set Dr, Trained on=Raw book content2026.01 | 28.92 | 33.62 | 80.28 | 59.47 | 3.58 | |
| UNDIALTrained on=Raw book content2026.01 | 28.2 | 29.09 | 73.92 | 58.08 | 1.08 | |
| GA + KLForget set=QA samples (Df_QA), KL-divergence regularization=Retention set Dr2026.01 | 27.44 | 36.87 | 84.95 | 60.62 | 7.63 | |
| GA + KLKL-divergence regularization=Retention set Dr, Trained on=Raw book content2026.01 | 27.2 | 38.29 | 78.67 | 60.18 | -0.27 | |
| NPOTrained on=Raw book content2026.01 | 24.18 | 26.83 | 69.69 | 54.79 | -0.16 | |
| NPO + KLForget set=QA samples (Df_QA), KL-divergence regularization=Retention set Dr2026.01 | 21.55 | 25.6 | 26.38 | 60.55 | -33.85 | |
| Distill from GA modelTeacher=GA-unlearned model, Scope=Retain-free / Forget-set behavior only2026.01 | 18.04 | 20.32 | 77.82 | 57.55 | 23.38 | |
| SimNPOTrained on=Raw book content2026.01 | 17.6 | 21.41 | 43.09 | 60.4 | -9.15 | |
| Distill from GA modelTeacher=GA-unlearned model, Scope=Forget and retain sets2026.01 | 17.4 | 19.54 | 42.86 | 54.53 | -13.18 | |
| SFRTrained on=Raw book content2026.01 | 10.45 | 13.17 | 71.93 | 58.28 | 32.96 | |
| DUETForget set=D_query_f U Dr2026.01 | 4.27 | 5.98 | 78.33 | 61.45 | 55.9 | |
| FLATTrained on=Raw book content2026.01 | 0.47 | 0.64 | 58.33 | 58.92 | 42.51 | |
| GATrained on=Raw book content2026.01 | 0 | 0 | 0 | 24.87 | -48.76 | |
| GAForget set=QA samples (Df_QA)2026.01 | 0 | 0 | 75.8 | 36.45 | 38.62 |