ResearchBenchmarksQuestion Answering on MT-Bench speech-interaction Turn one and two averagedFollow7.08Reasoning ScoreUnmute1.08962.64484.25.7552Sep 26, 2025Evaluation ResultsMethodMethodLinksReasoning ScoreSTEM ScoreHumanities ScoreAverage ScoreUnmuteBack-end LLM=GPT-4.1,...Back-end LLM=GPT-4.1, latency=2.12025.097.088.357.677.7KAMEBack-end LLM=Claude-op...Back-end LLM=Claude-opus-4.1, latency=0.02025.095.726.536.436.23KAMEBack-end LLM=GPT-4.1,...Back-end LLM=GPT-4.1, latency=0.02025.095.446.657.196.43Moshilatency=0.0latency=0.02025.091.321.942.882.05