visual representations
Methods to encode and understand visual data, such as images or videos, that enable AI models to interpret and manipulate visual information effectively. This includes feature extraction and various techniques to represent visual features.
- FORLA: Federated Object-centric Representation Learning with Slot Attention
- Glance2Gaze: Efficient Vision-Language Models from Glance Fusion to Gaze Compression
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Multimodal Tabular Reasoning with Privileged Structured Information
- Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
- Show-o2: Improved Native Unified Multimodal Models
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
- Token Bottleneck: One Token to Remember Dynamics
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models