Dialogue Generation on Norm-grounded Dialogue English (test)
65Win Rate vs. NormDialGPT-4o-mini
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4o-miniModel=GPT-4o-mini2025.09 | 65 | 65 | |
| LLaMA-3-8BModel=LLaMA-3-8B2025.09 | 65 | 59 | |
| Qwen-2.5-32BModel=Qwen-2.5-32B2025.09 | 56 | 62 | |
| Qwen-2.5-14BModel=Qwen-2.5-14B2025.09 | 51 | 61 |