Validating Response Generation on EmoValidBench Japanese
90.16BERT ScoreGPT 4.1 Nano
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT 4.1 NanoStrategy=3-shot2026.06 | 90.16 | 23.16 | 5 | 20.33 | 57.7 | 77.6 | 60.83 | 47.83 | 5.61 | 4.99 | 6.95 | 5.63 | 7 | 74.24 | |
| Llama 3.1 8bStrategy=3-shot2026.06 | 89.55 | 19.92 | 5.79 | 24.97 | 54.67 | 74.27 | 61.1 | 47.18 | 4.37 | 3.77 | 6.33 | 4.42 | 6.99 | 64.63 | |
| GPT 4.1 NanoStrategy=Zero-shot2026.06 | 89.54 | 19.96 | 4.56 | 18.4 | 53.34 | 76.77 | 61.62 | 46.31 | 5.49 | 4.9 | 6.94 | 5.58 | 7 | 73.42 | |
| Llama 3.1 8bStrategy=LoRA2026.06 | 89.34 | 16.56 | 3.39 | 15.02 | 56.92 | 75.38 | 58.84 | 45.06 | 3.03 | 2.7 | 5.66 | 3.29 | 6.95 | 56.32 | |
| Llama 3.1 8bStrategy=Zero-shot2026.06 | 89.15 | 18.24 | 5.23 | 23.67 | 54.73 | 72.07 | 61.38 | 46.35 | 3.93 | 3.35 | 6.38 | 4.04 | 7 | 59.79 | |
| GPT 4.1 NanoStrategy=CoT2026.06 | 88.14 | 14.38 | 7 | 27.74 | 49.9 | 78.09 | 61.35 | 46.66 | 5.72 | 5.16 | 6.93 | 5.73 | 7 | 74.3 |