text-to-video diffusion models

Generative models that create video content based on textual input, leveraging diffusion processes to progressively construct video outputs, representing an advanced intersection of natural language processing and computer vision.

4 papers