clip

CLIP (Contrastive Language-Image Pretraining) is a model developed by OpenAI that learns to understand images and text together, allowing for tasks like zero-shot image classification by aligning textual and visual semantics.

13 papers