Speech Quality Assessment on SPEAKBENCH
76Content ScoreHuman–human agreement
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Human–human agreementJudge=Human2026.01 | 76 | 60 | 82 | 60 | |
| TRACEBackbone=Gemini 2.5 Flash, Judge=Tree-based fusion2026.01 | 63.2 | 50.4 | 39.6 | — | |
| Audio JudgeJudge=Audio-only2026.01 | 62.5 | 45.6 | 21.4 | — | |
| LLM JudgeJudge=Transcript-only2026.01 | 60.4 | 39.8 | 29.8 | — | |
| Random GuessJudge=Random2026.01 | 25 | 25 | 25 | 25 |