Medical Terminology Error Detection on ChatCoach (test)
76.6BLEU-2Human
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human2024.02 | 76.6 | 6 | 90.5 | |
| Instruction-TuningCategory=Training-based, Backbone=Chinese_Alpaca2_LORA_13B, Fine-tuning method=QLORA2024.02 | 39.8 | 3 | 77.8 | |
| GCoTCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 34.2 | 3.7 | 72.4 | |
| Zero-shot CoTCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 27.6 | 1.9 | 69 | |
| Instruction PromptingCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 27.4 | 3.3 | 67.6 | |
| Vanilla CoTCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 17.7 | 2.7 | 64.1 |