text-to-video
- DisMo: Disentangled Motion Representations for Open-World Motion Transfer
- Rare Text Semantics Were Always There in Your Diffusion Transformer
- UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
- VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation