Video Motion Transfer on Video Motion Transfer Dataset 50 videos 1.0 (test)
35Text SimilarityDiTFlow
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DiTFlowBackbone=CogVideoX, Param=2B, Training Protocol=Training-free, Inference Time (s)=349, GPU Memory (G)=63.52026.03 | 35 | 69.1 | 98.3 | |
| FlowMotionBackbone=Wan2.1, Param=1.3B, Training Protocol=Training-free, Inference Time (s)=213, GPU Memory (G)=19.32026.03 | 34.7 | 85 | 98.6 | |
| DeTBackbone=CogVideoX, Param=2B, Training Protocol=Training-based, Training Time (s)=2760, Inference Time (s)=133, GPU Memory (G)=20.02026.03 | 34 | 81.2 | 98 | |
| MOFTBackbone=AnimateDiff, Param=1.3B, Training Protocol=Training-free, Inference Time (s)=576, GPU Memory (G)=75.02026.03 | 33.8 | 58.2 | 97.3 | |
| MotionDirectorBackbone=ZeroScope, Param=0.7B, Training Protocol=Training-based, Training Time (s)=1662, Inference Time (s)=140, GPU Memory (G)=28.02026.03 | 33.5 | 80.1 | 96.9 | |
| MotionCloneBackbone=AnimateDiff, Param=1.3B, Training Protocol=Training-free, Inference Time (s)=804, GPU Memory (G)=51.52026.03 | 33.2 | 78.6 | 94 | |
| MotionInversionBackbone=ZeroScope, Param=0.7B, Training Protocol=Training-based, Training Time (s)=1170, Inference Time (s)=115, GPU Memory (G)=24.02026.03 | 32.8 | 83.9 | 97 | |
| LoRA TuningBackbone=Wan2.1, Param=1.3B, Training Protocol=Training-based, Training Time (s)=8100, Inference Time (s)=135, GPU Memory (G)=25.02026.03 | 32.7 | 78.2 | 97.7 | |
| SMMBackbone=ZeroScope, Param=0.7B, Training Protocol=Training-free, Inference Time (s)=1839, GPU Memory (G)=89.42026.03 | 32.2 | 76.2 | 95.8 |