Spoken Language Assessment on Speak & Improve (S&I) Corpus 2025 (eval)
0.36RMSEPhi-4-MTL-APP
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Phi-4-MTL-APPFamily=MLLM2026.06 | 0.36 | 0.827 | 85.7 | 99 | |
| Whisper-only + SALR + LOPAFamily=Lightweight2026.06 | 0.361 | 0.828 | 83.3 | 99 | |
| Phi-4-MTLFamily=MLLM2026.06 | 0.362 | 0.825 | 85.7 | 99 | |
| Perezoso (Whisper + BERT + handcrafted features)Family=Lightweight2026.06 | 0.364 | 0.826 | 83 | 99.7 | |
| Phi-4-STGFamily=MLLM2026.06 | 0.375 | 0.82 | 81.7 | 99.3 | |
| APP (Whisper last-layer)Family=Lightweight2026.06 | 0.383 | 0.805 | 81.7 | 99 | |
| W2V (wav2vec2 end-to-end grader)Family=Lightweight2026.06 | 0.394 | 0.79 | 81.3 | 99.3 | |
| Phi-4-CTGFamily=MLLM2026.06 | 0.412 | 0.796 | 74.7 | 98 | |
| BERT (cascaded ASR → BERT baseline)Family=Lightweight2026.06 | 0.445 | 0.727 | 76 | 96.3 |