multi-modal large language models

Language models that can handle and integrate multiple types of data, such as text, images, and audio. These models typically leverage techniques from natural language processing and computer vision to understand and generate content across modalities.

17 papers