AI Tutoring Dialogue Alignment on Socratic Mind
70.47AccuracyMODPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MODPOBackbone=Qwen2.5-7B-Instruct2025.10 | 70.47 | 36 | |
| Single-Head DPOBackbone=Qwen2.5-7B-Instruct, Description=DPO applied to a single primary objective by pooling preference data2025.10 | 70.4 | 44.6 | |
| MAH-DPO Acc HeadBackbone=Qwen2.5-7B-Instruct, Head=Accuracy-specialized head2025.10 | 70.07 | 44.47 | |
| MAH-DPO Eng HeadBackbone=Qwen2.5-7B-Instruct, Head=Engagement-specialized head2025.10 | 69.53 | 44.8 | |
| MAH-DPO EnsembleBackbone=Qwen2.5-7B-Instruct, Ensemble=Equal weights2025.10 | 68.93 | 45.13 | |
| SFTBackbone=Qwen2.5-7B-Instruct, Description=Supervised fine-tuning on preferred responses2025.10 | 67.93 | 34.73 | |
| BaseBackbone=Qwen2.5-7B-Instruct, Description=LLM without post-training2025.10 | 65.6 | 32.2 |