mllms

Multimodal large language models that process and generate content across different modalities (e.g., text, image, audio). These models enable richer interactions and understanding by integrating multiple forms of data.

11 papers