Fine-tuning for knowledge acquisition and abstention preservation on TOFU
100FT ScoreFull-FT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Full-FTBase Model=Llama3-8B-Instruct2025.06 | 100 | 0 | 0 | |
| LoRABase Model=Llama3-8B-Instruct2025.06 | 100 | 21.5 | 12.7 | |
| Full-FTBase Model=Qwen2.5-7B-Instruct2025.06 | 100 | 0 | 0 | |
| LoRABase Model=Qwen2.5-7B-Instruct2025.06 | 100 | 23.6 | 15.2 | |
| EWCBase Model=Qwen2.5-7B-Instruct2025.06 | 100 | 24.6 | 14.7 | |
| CLoRABase Model=Qwen2.5-7B-Instruct2025.06 | 100 | 35.1 | 24.6 | |
| R-tuningBase Model=Qwen2.5-7B-Instruct2025.06 | 100 | 28.8 | 18.9 | |
| SEATBase Model=Qwen2.5-7B-Instruct2025.06 | 99.9 | 90.9 | 99.4 | |
| R-tuningBase Model=Llama3-8B-Instruct2025.06 | 99.8 | 2.6 | 2.1 | |
| Exp. ReplayBase Model=Llama3-8B-Instruct2025.06 | 99.7 | 37.7 | 48.7 | |
| Exp. ReplayBase Model=Qwen2.5-7B-Instruct2025.06 | 99.7 | 76.4 | 73.3 | |
| SEATBase Model=Llama3-8B-Instruct2025.06 | 98.7 | 96.5 | 97.7 | |
| EWCBase Model=Llama3-8B-Instruct2025.06 | 98.1 | 8.9 | 6.8 | |
| CLoRABase Model=Llama3-8B-Instruct2025.06 | 97.5 | 6.8 | 16.2 |