large vision-language models

Models that integrate visual and linguistic data to perform tasks that require understanding both modalities, such as image captioning and visual question answering. These models leverage large datasets of paired images and text.

28 papers