Criterion-Conditional In-Context Learning on CC-Bench Medical Diagnosis 1.0 (test)
14.26CS ScoreQwen2.5-VL-32B-Instruct
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen2.5-VL-32B-InstructModel Category=Open-Source Models, Support Set=3-shot2026.06 | 14.26 | 35.1 | 24.81 | |
| Qwen3-VL-8B-InstructModel Category=Open-Source Models, Support Set=3-shot2026.06 | 13.23 | 53.47 | 33.61 | |
| InternVL3.5-8BModel Category=Open-Source Models, Support Set=3-shot2026.06 | 9.85 | 41.04 | 25.65 | |
| GPT-4oModel Category=Proprietary API Models, Support Set=3-shot2026.06 | 9.76 | 34.19 | 22.13 | |
| Qwen2.5-VL-7B-InstructModel Category=Open-Source Models, Support Set=3-shot2026.06 | 8.54 | 31.35 | 20.09 | |
| LLaVA-OneVision-1.5-8B-InstructModel Category=Open-Source Models, Support Set=3-shot2026.06 | 7.5 | 29.43 | 18.61 | |
| Gemini 2.5 ProModel Category=Proprietary API Models, Support Set=3-shot2026.06 | 3.38 | 60.69 | 32.41 |