semantic consistency
The degree to which an AI-generated output maintains coherence and logical relevance with respect to the input context and semantics.
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- Aligning Text to Image in Diffusion Models is Easier Than You Think
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and Modeling
- Generative Pre-trained Autoregressive Diffusion Transformer
- HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
- Learning from Disjoint Views: A Contrastive Prototype Matching Network for Fully Incomplete Multi-View Clustering
- OPTFM: A Scalable Multi-View Graph Transformer for Hierarchical Pre-Training in Combinatorial Optimization
- Pancakes: Consistent Multi-Protocol Image Segmentation Across Biomedical Domains
- Pragmatic Heterogeneous Collaborative Perception via Generative Communication Mechanism
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
- T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks
- TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot Learning
- Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
- Vanish into Thin Air: Cross-prompt Universal Adversarial Attacks for SAM2