vision encoders
Vision encoders are neural network architectures designed to process and extract features from visual input, often used in computer vision tasks like image recognition, object detection, and segmentation.
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Perception Encoder: The best visual embeddings are not at the output of the network
- Perception Encoder: The best visual embeddings are not at the output of the network