Spoken Medical Question Answering on MMLU (test)
95.4ClinicalMedSpeak
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MedSpeakTranscription Source=Whisper generated transcripts, Fine-tuning Status=Fine-tuned, Knowledge Graph Context Integration=Budgeted KG context2026.02 | 95.4 | 93.3 | 95.6 | 95.8 | 97.8 | |
| Fine-Tuned LLM + GT (FT-LLM)Transcription Source=Ground truth texts, Fine-tuning Status=Fully fine-tuned, Knowledge Graph Context Integration=No KG-context2026.02 | 94.3 | 94.1 | 92.7 | 91 | 94.9 | |
| Fine-Tuned LLM + Whisper (FT+Whisp)Transcription Source=Whisper generated transcripts, Fine-tuning Status=Fine-tuned, Knowledge Graph Context Integration=No KG-context2026.02 | 85.6 | 85.2 | 84.2 | 85.5 | 88.6 | |
| Zero-shot GT (Zero-Shot)Transcription Source=Ground truth transcripts, Fine-tuning Status=No fine-tuning, Knowledge Graph Context Integration=No KG-context2026.02 | 66.3 | 64.4 | 67.5 | 48.8 | 76.1 | |
| Zero-shot ASR (ZS-ASR)Transcription Source=Whisper generated transcript, Fine-tuning Status=No fine-tuning, Knowledge Graph Context Integration=No KG-context2026.02 | 62.2 | 57 | 62.1 | 47.6 | 74.6 |