Text-to-Video Generation on VBench (aggregated)
73.89Semantic ScoreSelf Forcing + TokenTrim
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Self Forcing + TokenTrimInference Strategy=Self Forcing, Inference-time technique=TokenTrim, Backbone=Wan2.1-1.3B2026.01 | 73.89 | 89.79 | 81.84 | |
| Rolling Forcing + TokenTrimInference Strategy=Rolling Forcing, Inference-time technique=TokenTrim, Backbone=Wan2.1-1.3B2026.01 | 72.05 | 87.3 | 79.67 | |
| Rolling Forcing + FlowMoInference Strategy=Rolling Forcing, Inference-time technique=FlowMo, Backbone=Wan2.1-1.3B2026.01 | 69.53 | 82.09 | 75.81 | |
| Self ForcingInference Strategy=Self Forcing, Backbone=Wan2.1-1.3B2026.01 | 68.98 | 82.89 | 75.93 | |
| Rolling ForcingInference Strategy=Rolling Forcing, Backbone=Wan2.1-1.3B2026.01 | 68.52 | 81.72 | 75.12 | |
| Self Forcing + FlowMoInference Strategy=Self Forcing, Inference-time technique=FlowMo, Backbone=Wan2.1-1.3B2026.01 | 68.25 | 83.85 | 76.05 |