Task-oriented Dialogue on FewShotWeather unseen structures (test)
62.44BLEUFull Data Baseline
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Full Data BaselineTrain split=16,8162021.10 | 62.44 | 65.47 | |
| BLEURT self-trainingPseudo-response selection strategy=BLEURT, Train split=1shot-10002021.10 | 57.11 | 62.48 | |
| BLEURT self-trainingPseudo-response selection strategy=BLEURT, Train split=1shot-5002021.10 | 56.12 | 55.3 | |
| BLEURT self-trainingPseudo-response selection strategy=BLEURT, Train split=1shot-7502021.10 | 55.21 | 58.89 | |
| Vanilla self-trainingPseudo-response selection strategy=Vanilla, Train split=1shot-10002021.10 | 55.04 | 60.09 | |
| T5-smallPseudo-response selection strategy=None, Train split=1shot-7502021.10 | 54.49 | 54.02 | |
| Vanilla self-trainingPseudo-response selection strategy=Vanilla, Train split=1shot-7502021.10 | 54.32 | 54.19 | |
| Vanilla self-trainingPseudo-response selection strategy=Vanilla, Train split=1shot-5002021.10 | 54.27 | 49.91 | |
| T5-smallPseudo-response selection strategy=None, Train split=1shot-10002021.10 | 53.97 | 55.64 | |
| T5-smallPseudo-response selection strategy=None, Train split=1shot-5002021.10 | 53.62 | 46.58 | |
| BLEURT self-trainingPseudo-response selection strategy=BLEURT, Train split=1shot-2502021.10 | 52.34 | 43.68 | |
| Vanilla self-trainingPseudo-response selection strategy=Vanilla, Train split=1shot-2502021.10 | 51.87 | 31.37 | |
| T5-smallPseudo-response selection strategy=None, Train split=1shot-2502021.10 | 50.4 | 29.83 |