LLM Preference Alignment Evaluation on Math-DPO
78Preference Rate (spec vs ctl)Spec Learning
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Spec LearningProposer (P)=Kimi K2.6, Judge Model=GLM-5.12026.06 | 78 | 75 | 79 |
| Method | Links | |||
|---|---|---|---|---|
| Spec LearningProposer (P)=Kimi K2.6, Judge Model=GLM-5.12026.06 | 78 | 75 | 79 |