LLM Preference Alignment Evaluation on Truthy-DPO
64Spec vs ctl Preference ScoreSpec Learning
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Spec LearningProposer (P)=DeepSeek V4 Flash, Judge Model=GLM-5.12026.06 | 64 | 80 | 61 |
| Method | Links | |||
|---|---|---|---|---|
| Spec LearningProposer (P)=DeepSeek V4 Flash, Judge Model=GLM-5.12026.06 | 64 | 80 | 61 |