Multi-label Classification on Liver CT reports (test)
89.7CystGPT-4 (ICL)
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4 (ICL)In-Context Learning=three random samples from the training set2026.05 | 89.7 | 88.8 | 76.3 | 92.5 | 92 | 78.5 | 94.9 | 87.5 | 87.4 | |
| PromptRad+AutoTThresholding=Automatic, Evaluation=Average from five runs2026.05 | 89.5 | 90.8 | 78.4 | 91 | 97.3 | 84.7 | 92.4 | 89.2 | 89.4 | |
| GPT-4mode=Zero-shot2026.05 | 86.1 | 91.5 | 79.1 | 96.6 | 98.9 | 73.8 | 95.1 | 88.7 | 88.7 | |
| PromptRadEvaluation=Average from five runs2026.05 | 78 | 89.1 | 76.9 | 86.2 | 95.7 | 71.9 | 88.4 | 83.7 | 84.1 | |
| MetaMap2026.05 | 77.6 | 86.9 | 48.6 | 53.5 | 95.7 | 27.5 | 84.6 | 67.8 | 69.1 | |
| NegBio2026.05 | 77.6 | 86.9 | 81.4 | 82.7 | 95.7 | 27.5 | 84.6 | 76.6 | 79.2 | |
| PubMedBERT+MetaMapIntegration=MetaMap, Evaluation=Average from five runs2026.05 | 68.1 | 83.9 | 64.5 | 68 | 84.5 | 47.1 | 70.1 | 69.5 | 70.3 | |
| PubMedBERT+NegBioIntegration=NegBio, Evaluation=Average from five runs2026.05 | 68.1 | 83.9 | 76.8 | 81 | 84.5 | 47.1 | 70.1 | 73.1 | 74.1 | |
| Label Match2026.05 | 67.5 | 88.5 | 0 | 87.7 | 0 | 48.6 | 98.3 | 55.8 | 62.3 | |
| PubMedBERTEvaluation=Average from five runs2026.05 | 53.7 | 80.4 | 70.4 | 64.5 | 48.3 | 54.9 | 37.7 | 58.6 | 60.9 |