Text-to-Video Generation on VideoPhy-2
28.86SA ScoreCogVideoX-5B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CogVideoX-5BDetailed Prompts=false2026.04 | 28.86 | 68.42 | |
| MoAlign-2B (paper)Scale=2B, Implementation source=paper2026.05 | 28.8 | 75 | |
| PHANTOM-5BScale=5B2026.05 | 27.8 | 71.7 | |
| PhantomDetailed Prompts=false2026.04 | 27.75 | 71.74 | |
| PhantomDetailed Prompts=true2026.04 | 27.75 | 71.74 | |
| Cosmos-Diffusion-7BDetailed Prompts=false2026.04 | 26.32 | 54.19 | |
| VideoCrafter2Detailed Prompts=false2026.04 | 25.89 | 55.67 | |
| LaMo-2BScale=2B2026.05 | 25.4 | 75.4 | |
| MoAlign-2B (reimpl.)Scale=2B, Implementation source=reimpl.2026.05 | 24.6 | 73.1 | |
| Wan2.2-TI2V-5BDetailed Prompts=false2026.04 | 24.53 | 69.2 | |
| Wan2.2-TI2V-5BDetailed Prompts=true2026.04 | 24.53 | 69.2 | |
| Wan2.1-T2V-14B+VPTModel Backbone=Wan2.1, Parameter Scale=14B, Adaptation Method=VPT2026.07 | 23.3 | 59.9 | |
| Wan2.1-T2V-1.3B+VPTModel Backbone=Wan2.1, Parameter Scale=1.3B, Adaptation Method=VPT2026.07 | 22.5 | 55.1 | |
| Wan2.1-T2V-14BModel Backbone=Wan2.1, Parameter Scale=14B, Adaptation Method=None2026.07 | 21.9 | 52.9 | |
| VideoREPADetailed Prompts=false2026.04 | 21.02 | 72.54 | |
| VideoREPADetailed Prompts=true2026.04 | 21.02 | 72.54 | |
| CogVideoX-2BScale=2B2026.05 | 21 | 68 | |
| VideoREPA-2BScale=2B2026.05 | 21 | 72.5 | |
| Wan2.1-T2V-14B (Full Fine-tune)Model Backbone=Wan2.1, Parameter Scale=14B, Adaptation Method=Full Fine-tune2026.07 | 20.7 | 54 | |
| Wan2.1-T2V-1.3B+VideoJAMModel Backbone=Wan2.1, Parameter Scale=1.3B, Adaptation Method=VideoJAM2026.07 | 20.6 | 54 | |
| Wan2.1-T2V-1.3BModel Backbone=Wan2.1, Parameter Scale=1.3B, Adaptation Method=None2026.07 | 19.3 | 53.7 | |
| Wan2.1-T2V-1.3B (Full Fine-tune)Model Backbone=Wan2.1, Parameter Scale=1.3B, Adaptation Method=Full Fine-tune2026.07 | 18.9 | 53.6 |