Multimodal Healthcare Agent Performance on OpenHospital OOD
79.7SWAGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Zero-shot=true2026.03 | 79.7 | 34.8 | |
| CarePilotBackbone=Qwen 3 VL-8B2026.03 | 79.27 | 38.18 | |
| CarePilotBackbone=Qwen-VL 2.5-7B2026.03 | 77.93 | 36.4 | |
| Qwen3 VLModel size=235B, Zero-shot=true2026.03 | 75.18 | 25.46 | |
| GPT-4oZero-shot=true2026.03 | 74.63 | 25.48 | |
| Gemini 2.5 ProZero-shot=true2026.03 | 73.9 | 18.87 | |
| Llama 4 MaverickZero-shot=true2026.03 | 73.71 | 27.27 | |
| Nemotron VLModel size=12B, Zero-shot=true2026.03 | 72.9 | 18.18 | |
| Llama 4 ScoutZero-shot=true2026.03 | 72.2 | 20.75 | |
| Qwen2.5 VLModel size=32B, Zero-shot=true2026.03 | 71.74 | 12.72 | |
| Llama 3.2Model size=11B, Zero-shot=true2026.03 | 70.76 | 16.36 | |
| Mistral 3.2 VLModel size=24B, Zero-shot=true2026.03 | 69.63 | 1.82 |