Text Generation on MIMIC-RRS 100 samples (test)
4.414MedHelm LLM-jury scoreCAPT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CAPTComponents=Mnew + Mold-clin + Mold2026.01 | 4.414 | 56.67 | |
| Qwen3-30BModel Category=New-gen general (Mnew)2026.01 | 4.394 | 65 | |
| UniTEComponents=Mnew + Mold-clin2026.01 | 3.903 | 46.67 | |
| Proxy TuningComponents=Mold-L + Mold-clin + Mold2026.01 | 3.898 | 34.17 | |
| MeLLaMA-13B-chatModel Category=Old-gen clinical (Mold-clin)2026.01 | 3.379 | 35 | |
| LLaMA-2-70B-chatModel Category=Old-gen large general (Mold-L)2026.01 | 2.453 | 57.5 |