transformer model

A type of deep learning architecture that utilizes self-attention mechanisms, allowing for parallel processing of input data, crucial for handling sequential data in tasks like language modeling.

5 papers