image generation
The process of creating new images using generative models, which learn the underlying distribution of a dataset. Techniques can include GANs, VAEs, and diffusion models that leverage learned features to produce novel content.
- CAT: Content-Adaptive Image Tokenization
- Color Conditional Generation with Sliced Wasserstein Guidance
- Composition and Alignment of Diffusion Models using Constrained Learning
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
- Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- Differentiable Generalized Sliced Wasserstein Plans
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers
- Enhancing Consistency of Flow-Based Image Editing through Kalman Control
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
- Fourier Token Merging: Understanding and Capitalizing Frequency Domain for Efficient Image Generation
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- Inference-Time Personalized Alignment with a Few User Preference Queries
- KLASS: KL-Guided Fast Inference in Masked Diffusion Models
- LMFusion: Adapting Pretrained Language Models for Multimodal Generation
- Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation
- MLEP: Multi-granularity Local Entropy Patterns for Generalized AI-generated Image Detection
- More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
- NFIG: Multi-Scale Autoregressive Image Generation via Frequency Ordering
- ObCLIP: Oblivious CLoud-Device Hybrid Image Generation with Privacy Preservation
- PID-controlled Langevin Dynamics for Faster Sampling on Generative Models
- ReDi: Rectified Discrete Flow
- Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation
- RrED: Black-box Unsupervised Domain Adaptation via Rectifying-reasoning Errors of Diffusion
- SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score
- Show-o2: Improved Native Unified Multimodal Models
- Simple Distillation for One-Step Diffusion Models
- SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image Editing
- The Promise of RL for Autoregressive Image Editing
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Towards Understanding the Mechanisms of Classifier-Free Guidance
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement Learning
- When Worse is Better: Navigating the Compression Generation Trade-off In Visual Tokenization