video diffusion models
A class of generative models specifically designed for synthesis or analysis of video data, utilizing diffusion processes to iteratively refine video frames, aiming for high-quality output and temporal coherence.
- $\text{S}^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation
- EchoShot: Multi-Shot Portrait Video Generation
- Emergent Temporal Correspondences from Video Diffusion Transformers
- Event-Guided Consistent Video Enhancement with Modality-Adaptive Diffusion Pipeline
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
- MagCache: Fast Video Generation with Magnitude-Aware Cache
- MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
- UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting
- Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision