LLM-as-a-Judge Alignment on Personalized Response Quality human-annotated set
4.019Average RatingLlama-3.3-70B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama-3.3-70B2026.06 | 4.019 | 0.376 | |
| GPT-5.42026.06 | 3.523 | 0.312 | |
| Qwen3.5-27B2026.06 | 3.487 | 0.182 | |
| Claude-S4.62026.06 | 3.428 | 0.362 | |
| Gemma-4-31B2026.06 | 3.409 | 0.111 | |
| Human2026.06 | 3.176 | — |