Audio-driven Video Generation on Custom evaluation dataset
5.62Sync-COmnihuman-1
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Omnihuman-1Model Scale=–, Model Category=Closed-Source Audio-driven2025.12 | 5.62 | 8.8 | 0.82 | 0.9 | 0.65 | — | |
| Our teacher modelModel Scale=14B, Model Category=Open-Source Audio-driven Bidirectional DiT2025.12 | 5.56 | 8.19 | 0.78 | 0.91 | 0.63 | — | |
| JoyAvatar-FlashModel Scale=1.3B, Model Category=Audio-driven Causal DiT2025.12 | 5.25 | 8.3 | 0.72 | 0.84 | 0.53 | 16.2 | |
| HeygenModel Scale=–, Model Category=Closed-Source Audio-driven2025.12 | 4.86 | 9.09 | 0.79 | 0.92 | 0.66 | — | |
| WanS2VModel Scale=14B, Model Category=Open-Source Audio-driven Bidirectional DiT2025.12 | 4.05 | 9.38 | 0.67 | 0.88 | 0.47 | — | |
| LiveAvatarModel Scale=14B, Model Category=Audio-driven Causal DiT2025.12 | 3.89 | 9.45 | 0.78 | 0.85 | 0.52 | 4.3 | |
| OmniAvatarModel Scale=1.3B, Model Category=Open-Source Audio-driven Bidirectional DiT2025.12 | 3.85 | 9.54 | 0.72 | 0.88 | 0.5 | — | |
| EchoMimicV3Model Scale=1.3B, Model Category=Open-Source Audio-driven Bidirectional DiT2025.12 | 2.49 | 10.78 | 0.72 | 0.9 | 0.51 | — | |
| StableAvatarModel Scale=1.3B, Model Category=Open-Source Audio-driven Bidirectional DiT2025.12 | 2.47 | 10.97 | 0.72 | 0.91 | 0.53 | — |