Human-Object Interaction Video Generation on GroundedInter Pose-Driven 1.0 (test)
28.47VLM-QAInteractAvatar
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| InteractAvatarInput Signal=Pose-Driven, Inference Mode=TAM2V2026.02 | 28.47 | 95.7 | 14.1 | 77.3 | 83.2 | 64.8 | 84.2 | 28.7 | 90.2 | 94 | 5.73 | |
| UniAnimate-DiTInput Signal=Pose-Driven2026.02 | 24.65 | 86.2 | 8.2 | 44.1 | 66.9 | 62.1 | 78.2 | 27.7 | 87.5 | 93.8 | — |