Validating Response Generation on 100 English and Japanese Utterance-Response Pairs (test)
5.48NaturalnessGround Truth
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Ground TruthType=Human responses2026.06 | 5.48 | 5.76 | 5.92 | 6.08 | 5.66 | 5.71 | 0.406 | |
| GPT-4.1 NanoPrompting=3-shot2026.06 | 5.42 | 5.67 | 5.59 | 5.85 | 5.43 | 5.47 | 0.322 | |
| Llama 3.1 8BPrompting=3-shot2026.06 | 5.27 | 5.67 | 5.35 | 5.98 | 5.53 | 5.33 | 0.453 |