Avatar Rendering and Audio-Visual Generation on System-specific evaluation sets
40FPSVASA-1
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| VASA-1Visual interaction scope=audio-driven talking face2026.06 | 40 | 0.17 | — | — | — | — | — | — | — | |
| OmniForcingVisual interaction scope=text-to-audio-video streaming generation2026.06 | 25 | — | — | — | 0.7 | — | — | — | — | |
| Wan-StreamerVisual interaction scope=text/audio/video perceptual dialogue with synchronized speech and video output2026.06 | 25 | — | — | — | — | — | — | 550 | 200 | |
| LiveTalkVisual interaction scope=multimodal interactive avatar video2026.06 | 24.82 | 0.0003 | — | — | — | — | — | — | — | |
| Hallo-LiveVisual interaction scope=text-driven joint audio-video avatar2026.06 | 20.38 | 0.0009 | — | — | — | — | — | — | — | |
| Avatar Forcing (Ki et al.)Visual interaction scope=interactive head-avatar reactions2026.06 | — | 0.5 | — | — | — | — | — | — | — | |
| AvatarForcing (Cui et al.)Visual interaction scope=one-step streaming talking avatar2026.06 | — | — | — | — | — | 34 | 0.51 | — | — | |
| StreamAvatarVisual interaction scope=streaming talking/listening avatar2026.06 | — | 0.0012 | 0.33 | — | — | — | — | — | — |