Long-form Response Generation for Clinical Questions on K-QA
81CompletenessGPT-4
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4Size=Unknown2024.03 | 81 | 92.5 | |
| MeerkatSize=70B2024.03 | 75.4 | 89.6 | |
| MeerkatSize=8B2024.03 | 72.2 | 90 | |
| GPT-3.5Size=175B2024.03 | 71.4 | 92 | |
| MeerkatSize=7B2024.03 | 70.3 | 89.6 | |
| ChatDoctorSize=7B2024.03 | 63 | 89.1 | |
| Mistral-InstructSize=7B2024.03 | 62.4 | 88.1 | |
| PMC-LLAMASize=13B2024.03 | 49.8 | 90 | |
| Med-AlpacaSize=13B2024.03 | 6.8 | — |