Human-motion video generation on Motion-X
1.549VA-MQVideoAlign-MQ
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| VideoAlign-MQTraining Strategy=VLM-reward trained2026.05 | 1.549 | 0.912 | 4.12 | 0.916 | 0.753 | 0.946 | 0.878 | |
| PhyMotionTraining Strategy=Structured 3D reward training, Base Model=FastWan 1.3B2026.05 | 1.318 | 0.687 | 4.39 | 0.982 | 0.805 | 0.952 | 0.913 | |
| PhyMotionTraining Strategy=Structured 3D reward training, Base Model=Causal Forcing 1.3B2026.05 | 1.313 | 0.72 | 4.29 | 0.982 | 0.809 | 0.988 | 0.902 | |
| FastWan 1.3BModel Scale=1.3B2026.05 | 1.307 | 0.68 | 3.9 | 0.908 | 0.734 | 0.916 | 0.853 | |
| Causal Forcing 1.3BModel Scale=1.3B2026.05 | 1.241 | 0.575 | 4.06 | 0.927 | 0.739 | 0.948 | 0.871 | |
| Wan 1.3BModel Scale=1.3B2026.05 | 1.213 | 0.621 | 3.86 | 0.862 | 0.705 | 0.904 | 0.824 | |
| Wan2.2 14BModel Scale=14B2026.05 | 1.133 | 0.557 | 3.97 | 0.881 | 0.733 | 0.912 | 0.842 | |
| Wan2.2 5BModel Scale=5B2026.05 | 1.069 | 0.514 | 3.91 | 0.88 | 0.71 | 0.913 | 0.834 | |
| EchoMotion 5BModel Scale=5B2026.05 | 1.044 | 0.565 | 3.74 | 0.888 | 0.702 | 0.918 | 0.836 |