vision-language-action models

Models that integrate visual, linguistic, and action-oriented components to perform tasks like robotic control or interacting with environments based on visual and textual input.

7 papers