Clinical Diagnosis on MIMIC-CDM (test)
99.4Appendicitis Accuracygpt-4o
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| gpt-4oEvaluation Protocol=Inference-only API, Latent path modeling=No, Trajectory-level regularization=No2026.04 | 99.4 | 98.4 | 88.4 | 88.8 | 95.6 | — | — | — | 95.7 | |
| gemini-3-flashEvaluation Protocol=Inference-only API, Latent path modeling=No, Trajectory-level regularization=No2026.04 | 99.4 | 94.6 | 90.3 | 90.7 | 95.2 | — | — | — | 95.3 | |
| deepseek-chatEvaluation Protocol=Inference-only API, Latent path modeling=No, Trajectory-level regularization=No2026.04 | 99.4 | 97.6 | 92.3 | 92.5 | 96.6 | — | — | — | 96.6 | |
| claude-sonnet-4Evaluation Protocol=Inference-only API, Latent path modeling=No, Trajectory-level regularization=No2026.04 | 99.3 | 98.4 | 90.3 | 88.8 | 95.8 | — | — | — | 95.8 | |
| LDTLMethod description=Latent Diagnostic Trajectory Learning, Inference structure=Sequential2026.04 | 98.9 | 95.4 | 78.8 | 87.9 | 93.4 | — | — | — | 91.7 | |
| LDTLBackbone=Llama-3 8B, Evaluation Protocol=Parameter-efficient fine-tuning (LoRA), Latent path modeling=Yes, Trajectory-level regularization=Yes2026.04 | 98.9 | 95.4 | 78.8 | 87.9 | 93.4 | — | — | — | 91.7 | |
| SFT-allTest Access Strategy=All available patient information directly, Evaluation Protocol=Supervised fine-tuning2025.06 | 98.4 | 89.8 | 95.8 | 87.5 | 92.8 | 93.6 | 92.9 | 3,792.79 | — | |
| SFT-allEvaluation protocol=Fine-tuned (SFT), Information access=Complete patient information, Decision strategy=One-shot2026.04 | 97.9 | 93.1 | 90.4 | 90.7 | 94.2 | — | — | — | 95.8 | |
| SFT-allBackbone=Llama-3-8B, Evaluation Protocol=Supervised Fine-Tuning (SFT)2026.04 | 97.9 | 93.1 | 90.4 | 90.7 | 94.2 | — | — | — | 95.8 | |
| gpt-4o-miniEvaluation Protocol=Inference-only API, Latent path modeling=No, Trajectory-level regularization=No2026.04 | 94.7 | 91.5 | 82.6 | 86.1 | 90.5 | — | — | — | 90.6 | |
| LA-CDMEvaluation Protocol=Trained2025.06 | 93.1 | 83.6 | 75 | 73.5 | 81.3 | 84.1 | 81.3 | 1,295.61 | — | |
| LA-CDMInference structure=Reinforcement learning based, Decision strategy=Sequential agent framework2026.04 | 93.1 | 83.6 | 75 | 73.5 | 81.3 | — | — | — | 84.1 | |
| ReActEvaluation Protocol=Zero-shot2025.06 | 90.2 | 79.7 | 66.7 | 62.9 | 74.9 | 79.1 | 74.8 | 1,480.32 | — | |
| ReActInference structure=Prompting based agent, Decision strategy=Sequential agent framework2026.04 | 90.2 | 79.7 | 66.7 | 62.9 | 74.9 | — | — | — | 79.1 | |
| Planner random selectionDecision strategy=Random policy, Inference structure=Sequential2026.04 | 82.8 | 85.4 | 90.4 | 85.2 | 84.8 | — | — | — | 83.7 | |
| SM-DDPOData Modality=Tabular only, Action Space=Laboratory values only2025.06 | 74.3 | 0 | 15.6 | 58 | 37 | 45.4 | 31.8 | — | — | |
| All infoEvaluation protocol=Zero-shot (ZS), Information access=Complete patient information (Full-information), Decision strategy=One-shot2026.04 | 73.9 | 70 | 82.7 | 87.9 | 76.9 | — | — | — | 76.7 | |
| LA-CDM (ZS)Evaluation Protocol=Zero-shot, Status=Untrained version2025.06 | 73.5 | 55.1 | 72 | 57.5 | 64.5 | 65.3 | 64.5 | 1,521.73 | — | |
| Fixed info (history + two tests)Evaluation protocol=Zero-shot (ZS), Information access=History plus two of three test categories, Decision strategy=Static2026.04 | 60.9 | 60.7 | 78.8 | 80.5 | 67.2 | — | — | — | 68.2 | |
| Fixed info (history + one test)Evaluation protocol=Zero-shot (ZS), Information access=History plus exactly one test category, Decision strategy=Static2026.04 | 50.5 | 47.7 | 76.9 | 72.2 | 58.3 | — | — | — | 60.5 | |
| Fixed info (history)Evaluation protocol=Zero-shot (ZS), Information access=Patient history only, Decision strategy=Static2026.04 | 47.9 | 34.6 | 73.1 | 79.6 | 54.1 | — | — | — | 58.8 |