vision transformer

A vision transformer is a type of model architecture that applies transformer techniques—particularly self-attention—to image data, enhancing the model's ability to capture spatial relationships and context within images.

7 papers