Knowledge Preservation and Reasoning on MMLU
61.46MMLU ScoreBase Model (Llama3.2-3B)
Evaluation Results
| Method | Links | |
|---|---|---|
| Base Model (Llama3.2-3B)Backbone=Llama 3.2-3B-Instruct2026.01 | 61.46 | |
| DUETBackbone=Llama 3.2-3B-Instruct, Training Data Format=Dquery_f ∪ Dr2026.01 | 61.45 | |
| GA (DQA_f) + KL (Dr)Backbone=Llama 3.2-3B-Instruct, Training Data Format=QA samples, Regularization=KL-divergence2026.01 | 60.62 | |
| NPO (DQA_f) + KL (Dr)Backbone=Llama 3.2-3B-Instruct, Training Data Format=QA samples, Regularization=KL-divergence2026.01 | 60.55 | |
| NPO (DQA_f)Backbone=Llama 3.2-3B-Instruct, Training Data Format=QA samples2026.01 | 60.48 | |
| Refusal-TrainingBackbone=Llama 3.2-3B-Instruct, Training Data Format=DQR_f ∪ Dr2026.01 | 60.48 | |
| SimNPOBackbone=Llama 3.2-3B-Instruct2026.01 | 60.4 | |
| GA + KL (Dr)Backbone=Llama 3.2-3B-Instruct, Regularization=KL-divergence2026.01 | 60.18 | |
| NPO + KL (Dr)Backbone=Llama 3.2-3B-Instruct, Regularization=KL-divergence2026.01 | 59.47 | |
| FLATBackbone=Llama 3.2-3B-Instruct2026.01 | 58.92 | |
| NPOBackbone=Llama 3.2-3B-Instruct2026.01 | 54.79 | |
| GA (DQA_f)Backbone=Llama 3.2-3B-Instruct, Training Data Format=QA samples2026.01 | 36.45 | |
| GABackbone=Llama 3.2-3B-Instruct2026.01 | 24.87 |