Adaptive Care Policy Learning on ALPACA 1000 simulated patient rollouts
3.38Cumulative RewardPPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PPOtype=reinforcement learning agent2026.02 | 3.38 | 0.27 | -0.46 | |
| SACtype=reinforcement learning agent2026.02 | 0.18 | 0.01 | -0.76 | |
| A2Ctype=reinforcement learning agent2026.02 | -0.66 | -0.06 | -0.59 | |
| BDQtype=reinforcement learning agent2026.02 | -7.25 | -0.47 | -0.98 | |
| Cliniciantype=behavior cloned baseline2026.02 | -13.72 | -0.73 | -1.36 | |
| Heuristic2026.02 | -15.1 | -0.79 | -1.42 | |
| No Medicationtype=baseline2026.02 | -17.81 | -0.95 | -1.53 |