Satisfaction Prediction on human turn-level satisfaction annotations
0.3601Pearson CorrelationUser-memory evaluator
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| User-memory evaluatorType=User-aware LLM judge, Backbone=Qwen3-8B2026.05 | 0.3601 | 0.3716 | 0.3595 | |
| Nearest-history scoreType=User-history retrieval2026.05 | 0.3281 | 0.3529 | 0.2992 | |
| Prometheus-rubric judgeType=Generic LLM judge, Backbone=Qwen3-8B2026.05 | 0.2066 | 0.1435 | 0.2005 | |
| Few-shot judgeType=Generic LLM judge, Backbone=Qwen3-8B2026.05 | 0.1924 | 0.1578 | 0.1867 | |
| Zero-shot judgeType=Generic LLM judge, Backbone=Qwen3-8B2026.05 | 0.1865 | 0.1274 | 0.1753 | |
| Task-rubric judgeType=Generic LLM judge, Backbone=Qwen3-8B2026.05 | 0.16 | 0.121 | 0.1511 | |
| RAG prompted scorerType=User-history retrieval2026.05 | 0.136 | 0.0776 | 0.0282 | |
| SPUR-style evaluatorType=Generic LLM judge, Backbone=Qwen3-8B2026.05 | 0.1314 | 0.1095 | 0.0842 | |
| BERT ordinal scorerType=Supervised scorer, Backbone=BERT2026.05 | 0.0363 | 0.0625 | 0.0322 | |
| BERT scorerType=Supervised scorer, Backbone=BERT2026.05 | 0.0126 | 0.0192 | 0.0125 |