Motivational Interviewing Response Generation on AnnoMI (test)
2.1MI-i (%)GPT3.5-BESTBASE
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT3.5-BESTBASEBackbone=GPT-3.5, Strategy=Best of four baselines (base model and three ICL variants)2024.03 | 2.1 | 44.8 | 1.53 | 54.3 | 82.8 | |
| GPT4-BESTBASEBackbone=GPT-4, Strategy=Best of four baselines (base model and three ICL variants)2024.03 | 1.4 | 11.8 | 1.06 | 45.8 | 71.8 | |
| GoldMode=Human Ground Truth2024.03 | 1.4 | 2.3 | 1.51 | 59 | 81.7 | |
| GPT3.5-DIIRBackbone=GPT-3.5, Strategy=DIIR (Discovery and Inference of In-context Strategies)2024.03 | 1.3 | 115 | 0.99 | 61.1 | 95.1 | |
| GPT4-DIIRBackbone=GPT-4, Strategy=DIIR (Discovery and Inference of In-context Strategies)2024.03 | 0.6 | 135 | 1.04 | 59.6 | 94.4 |