generative models
Generative models are a class of AI models that can learn to generate new data points similar to those in the training set. This paradigm includes techniques like Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) and has applications in image synthesis, text generation, and more.
- $i$MIND: Insightful Multi-subject Invariant Neural Decoding
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
- Active Target Discovery under Uninformative Priors: The Power of Permanent and Transient Memory
- Adaptive 3D Reconstruction via Diffusion Priors and Forward Curvature-Matching Likelihood Updates
- Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- AlignedGen: Aligning Style Across Generated Images
- All that structure matches does not glitter
- Ambient Proteins - Training Diffusion Models on Noisy Structures
- Amortized Sampling with Transferable Normalizing Flows
- Antidistillation Sampling
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval
- BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
- Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
- Beyond Scores: Proximal Diffusion Models
- Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness Conditions
- BikeBench: A Bicycle Design Benchmark for Generative Models with Objectives and Constraints
- COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
- CPO: Condition Preference Optimization for Controllable Image Generation
- Can We Infer Confidential Properties of Training Data from LLMs?
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
- Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
- Constrained Discrete Diffusion
- Controllable 3D Molecular Generation for Structure-Based Drug Design Through Bayesian Flow Networks and Gradient Integration
- DP²O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- Denoising Trajectory Biases for Zero-Shot AI-Generated Image Detection
- Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based Decoding
- Diffusion Classifiers Understand Compositionality, but Conditions Apply
- Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
- ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators
- ESCA: Enabling Seamless Codec Avatar Execution through Algorithm and Hardware Co-Optimization for Virtual Reality
- Efficient Multimodal Dataset Distillation via Generative Models
- Energy Matching: Unifying Flow Matching and Energy-Based Models for Generative Modeling
- Energy-based generator matching: A neural sampler for general state space
- EngiBench: A Framework for Data-Driven Engineering Design Research
- EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation
- Epistemic Uncertainty for Generated Image Detection
- Failure Prediction at Runtime for Generative Robot Policies
- FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies
- Flash Invariant Point Attention
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior Modeling
- Flow-Based Policy for Online Reinforcement Learning
- Follow the Energy, Find the Path: Riemannian Metrics from Energy-Based Models
- Forging Time Series with Language: A Large Language Model Approach to Synthetic Data Generation
- FracFace: Breaking The Visual Clues—Fractal-Based Privacy-Preserving Face Recognition
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language Models
- Generating Informative Samples for Risk-Averse Fine-Tuning of Downstream Tasks
- GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
- Graph Diffusion that can Insert and Delete
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- Improving the Generation and Evaluation of Synthetic Data for Downstream Medical Causal Inference
- Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
- Information Theoretic Learning for Diffusion Models with Warm Start
- Injecting Frame-Event Complementary Fusion into Diffusion for Optical Flow in Challenging Scenes
- Is Artificial Intelligence Generated Image Detection a Solved Problem?
- Joint Relational Database Generation via Graph-Conditional Diffusion Models
- Knowledge Distillation Detection for Open-weights Models
- Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental Learning
- LaM-SLidE: Latent Space Modeling of Spatial Dynamical Systems via Linked Entities
- Learning to Generate Human-Human-Object Interactions from Textual Descriptions
- Leveraging semantic similarity for experimentation with AI-generated treatments
- LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
- Localizing Knowledge in Diffusion Transformers
- LuxDiT: Lighting Estimation with Video Diffusion Transformer
- MLEP: Multi-granularity Local Entropy Patterns for Generalized AI-generated Image Detection
- Machine Unlearning in 3D Generation: A Perspective-Coherent Acceleration Framework
- MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
- Multitask Learning with Stochastic Interpolants
- Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards
- Neural-Driven Image Editing
- Noise Matters: Optimizing Matching Noise for Diffusion Classifiers
- OASIS: One-Shot Federated Graph Learning via Wasserstein Assisted Knowledge Integration
- OSTAR: Optimized Statistical Text-classifier with Adversarial Resistance
- On the Relation between Rectified Flows and Optimal Transport
- On the Sample Complexity Bounds of Bilevel Reinforcement Learning
- One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- Orientation Matters: Making 3D Generative Models Orientation-Aligned
- PID-controlled Langevin Dynamics for Faster Sampling on Generative Models
- PhysX-3D: Physical-Grounded 3D Asset Generation
- PocketSR: The Super-Resolution Expert in Your Pocket Mobiles
- Prior-Guided Flow Matching for Target-Aware Molecule Design with Learnable Atom Number
- ProteinConformers: Benchmark Dataset for Simulating Protein Conformational Landscape Diversity and Plausibility
- ROSE: Remove Objects with Side Effects in Videos
- Rare Text Semantics Were Always There in Your Diffusion Transformer
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- ReDi: Rectified Discrete Flow
- Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
- RiboFlow: Conditional De Novo RNA Co-Design via Synergistic Flow Matching
- Riemannian Consistency Model
- SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
- SGCD: Stain-Guided CycleDiffusion for Unsupervised Domain Adaptation of Histopathology Image Classification
- Scaling Diffusion Transformers Efficiently via $\mu$P
- SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
- Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion Models
- Space Group Equivariant Crystal Diffusion
- Sparse Image Synthesis via Joint Latent and RoI Flow
- Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations
- Steering Generative Models with Experimental Data for Protein Fitness Optimization
- TARFVAE: Efficient One-Step Generative Time Series Forecasting via TARFLOW based VAE
- Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
- Towards foundational LiDAR world models with efficient latent flow matching
- Training-Free Constrained Generation With Stable Diffusion Models
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- Unified all-atom molecule generation with neural fields
- Unveiling Concept Attribution in Diffusion Models
- UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
- Value Gradient Guidance for Flow Matching Alignment
- VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
- Watermarking Autoregressive Image Generation
- What's Producible May Not Be Reachable: Measuring the Steerability of Generative Models
- When Models Don’t Collapse: On the Consistency of Iterative MLE
- Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation
- un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP