Multi-task Language Understanding on MMLU (7-shot, test)
70.13Humanities AccuracyBase model (Qwen2.5-3B-Instruct)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Base model (Qwen2.5-3B-Instruct)Fine-tuning=None2026.02 | 70.13 | 75.33 | 71.77 | 64.06 | |
| ZS fine-tuning on KFine-tuning objective=Zero-shot (ZS), LoRA adapted parameters=K2026.02 | 69.06 | 75.01 | 74.07 | 67.02 | |
| ZS fine-tuning on QFine-tuning objective=Zero-shot (ZS), LoRA adapted parameters=Q2026.02 | 68.57 | 73.13 | 71.17 | 64.01 | |
| ZS+FS fine-tuning on V (Value-Matrix Fine-Tuning)Fine-tuning objective=Zero-shot + Few-shot (ZS+FS), LoRA adapted parameters=V2026.02 | 68.36 | 75.56 | 65.29 | 64.7 | |
| ZS fine-tuning on V (Value-Matrix Fine-Tuning)Fine-tuning objective=Zero-shot (ZS), LoRA adapted parameters=V2026.02 | 67.75 | 72.86 | 71.35 | 64.74 | |
| ZS fine-tuning on Q/K/VFine-tuning objective=Zero-shot (ZS), LoRA adapted parameters=Q/K/V2026.02 | 66.06 | 73.26 | 66.72 | 62.91 |