Medical terminology error correction on ChatCoach (test)
33.5BLEU-2Human
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human2024.02 | 33.5 | 3.6 | 84.1 | |
| Instruction-TuningCategory=Training-based, Backbone=Chinese_Alpaca2_LORA_13B, Fine-tuning method=QLORA2024.02 | 4 | 1.7 | 59.7 | |
| Zero-shot CoTCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 3 | 0.9 | 58.8 | |
| GCoTCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 1.6 | 2 | 65.4 | |
| Instruction PromptingCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 1.4 | 2.1 | 61.6 | |
| Vanilla CoTCategory=Prompting-based, Base Model=gpt-3.5-turbo2024.02 | 0.1 | 2.3 | 58.1 |