Task-Focused Dialogue on Multiwoz (TSE, Reward, BLEU)
0.7699TSE ScoreGOPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GOPOBackbone=Qwen3-14B2026.01 | 0.7699 | 6.55 | 0.177 | |
| DeepSeekModel Version=R12026.01 | 0.7609 | 6.46 | 0.138 | |
| GLMModel Version=4.72026.01 | 0.7598 | 6.13 | 0.148 | |
| GPTModel Version=5.22026.01 | 0.7561 | 5.83 | 0.092 | |
| GOPOBackbone=Qwen-7B-Chat2026.01 | 0.7543 | 6.38 | 0.172 | |
| PPO2026.01 | 0.7496 | 5.68 | 0.169 | |
| QwenModel Version=235B2026.01 | 0.7462 | 5.57 | 0.165 | |
| GeminiModel Version=2.52026.01 | 0.7451 | 6.16 | 0.129 | |
| Memento2026.01 | 0.7205 | 5.75 | 0.153 | |
| SFT2026.01 | 0.7015 | 4.92 | 0.149 | |
| Untrained2026.01 | 0.6127 | 4.53 | 0.075 |