vision foundation models
These are large pre-trained models designed for various vision tasks (e.g., object detection, image segmentation) that serve as a baseline for transfer learning or fine-tuning on specific applications, often leveraging vast quantities of unlabeled data for training.
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- Connecting Neural Models Latent Geometries with Relative Geodesic Representations
- DON’T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object Detection
- Emergent Temporal Correspondences from Video Diffusion Transformers
- FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed
- Object Concepts Emerge from Motion
- Online Segment Any 3D Thing as Instance Tracking
- Revisiting Semi-Supervised Learning in the Era of Foundation Models