Human-Object Interaction Video Generation on GroundedInter Text-Driven 1.0 (test)
29.07VLM-QAInteractAvatar
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| InteractAvatarInput Signal=Text-Driven, Inference Mode=TA2V2026.02 | 29.07 | 97.3 | 14.7 | 78.3 | 85 | 64.1 | 83.9 | 29 | 90.4 | 94.2 | 5.92 | |
| InteractAvatarInput Signal=Text-Driven, Inference Mode=T2MV2026.02 | 29.05 | 97.5 | 15 | 79.2 | 85.2 | 63.7 | 83.5 | 28.9 | 90.4 | 93.2 | — | |
| VACEInput Signal=Text-Driven2026.02 | 26.74 | 90.8 | 11.8 | 71.9 | 70.5 | 62.9 | 81.7 | 28.5 | 89.3 | 94 | — | |
| HuMoInput Signal=Text-Driven2026.02 | 24.12 | 91 | 10.1 | 38 | 49.1 | 56.6 | 72.6 | 28.1 | 88 | 93.5 | 5.15 |