Helpful response generation on Human A/B 100 randomly chosen instances (test)
62Human Preference ScoreMuffinGPT-3.5
Evaluation Results
| Method | Links | |
|---|---|---|
| MuffinGPT-3.5base_model=GPT-3.52024.01 | 62 | |
| GPT-3.5Setting=in-context learning baseline2024.01 | 8 |
| Method | Links | |
|---|---|---|
| MuffinGPT-3.5base_model=GPT-3.52024.01 | 62 | |
| GPT-3.5Setting=in-context learning baseline2024.01 | 8 |