vision-language model

A vision-language model is an AI architecture designed to process and relate information from both visual inputs (like images and videos) and linguistic inputs (such as text descriptions). These models learn to associate visual content with natural language, enabling tasks like image captioning and visual question answering.

20 papers