Human Evaluation
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Human Evaluation (test)
3.12Fluency Score
4
Human evaluation Content relevance
4.3Evidence Score
4
Human Evaluation Interactive Gaming
4.31Action Controllability
4
Human evaluation 20 prompts (test)
81Win Rate
4
Human Evaluation 14 clips of two-person conversations (test)
79.3Lip Sync Score
4
Human evaluation (test)
4.78Lip-sync Score
4
Human Evaluation
0.86Trustworthiness
4
Human Evaluation 15 scenes 1.0
4.8Stereo Effect (SE)
4
Human Evaluation CAD Generation (test)
0.84Semantics
4
Human Evaluation 50 videos
74.1Win Rate (%)
4
Human Evaluation Subjective Audio Assessment (test)
0.249Z-Score (OVL)
4
Human Evaluation (Beginners)
4OVL
4
Human Evaluation (Experts)
4.2OVL (Overall Likeness)
4
Human Evaluation V2A
3.7Audio Quality
4
Human Evaluation Dataset (test)
53.75Near Object Score
4
Human Evaluation (120 randomly sampled cases from HotpotQA, 2WikiMultiHopQA, MuSiQue, and DuReader)
52.75Accuracy
4
Human Evaluation 300 samples (random sample)
0.182Score 3 (%)
4
Human Evaluation 300 samples (random sample)
28.7Score 3
4
Human Evaluation 2-hop questions (test)
74Well-formed Rate (Yes)
4
Human Evaluation
2.73Self-contain Accuracy
3
Human Evaluation Rapport and UX
3.88Rapport Score R1
3
Human Evaluation Chinese N=10
85Win Rate
3
Human Evaluation Turkish tr N=10
95Win Rate
3
Human Evaluation Thai N=10
95Win Rate
3
Human Evaluation Russian N=10
90Win Rate
3