Supportive Response Generation on Reddit Peer-Support Dataset (Phase 2 Human Evaluation)
4.59ReadabilityGPT-5-nano
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5-nanoevaluation_type=Human evaluation (mean rating), scale=1-5 Likert scale2026.05 | 4.59 | 4.342 | 3.965 | 4.322 | 4.59 | 9.491 | — | — | — | |
| LLUMI-GM (DPO2)evaluation_type=Human evaluation (mean rating), scale=1-5 Likert scale, phase=Phase 2 evaluation2026.05 | 4.387 | 4.413 | 4.059 | 4.238 | 4.487 | — | 6.499 | -3.321 | 64.616 | |
| Online Community (OC)evaluation_type=Human evaluation (mean rating), scale=1-5 Likert scale2026.05 | 3.812 | 2.135 | 1.941 | 1.803 | 2.95 | — | — | — | — |