Text-to-Video Generation on HumanVid 500 real-world videos (curated evaluation set)
0.46LPIPSTokenMotion-T
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| TokenMotion-Tbackbone=CogVideoX-2B2025.04 | 0.46 | 82.36 | 361.03 | 4.43 | 0.27 | 5.34 | 45.24 | 2.5 | |
| CogVideoX-2Bstatus=original model without motion control2025.04 | 0.57 | 88.11 | 402.34 | — | — | — | — | — | |
| MotionCtrl2025.04 | 0.63 | 116.41 | 1,185.83 | 6.72 | 0.47 | 8.66 | 172.44 | 14.73 | |
| Direct-A-Video2025.04 | 0.67 | 132.08 | 941.05 | 3.85 | 0.24 | 5.06 | 173.26 | 10.68 | |
| MotionBooth2025.04 | 0.69 | 125.62 | 795.59 | 5.24 | 0.36 | 6.42 | 165.49 | 13.19 |