Dialogue Evaluation on ACUTE-Eval Human-Chat (test)
75EngagingnessBlenderBot
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BlenderBotParameters=2.7B, Reference Study=Roller et al. (2020)2021.05 | 75 | 65 | |
| BlenderBotParameters=2.7B, Evaluation Toolkit=LEGOEval (reproduction)2021.05 | 72 | 68 | |
| MeenaEvaluation Toolkit=LEGOEval (reproduction)2021.05 | 28 | 32 | |
| MeenaReference Study=Roller et al. (2020)2021.05 | 25 | 35 |