sparse attention

Sparse attention is a mechanism used in neural networks, particularly in transformers, which focuses on only a subset of input elements during training and inference, improving efficiency and reducing computational overhead while retaining performance.

10 papers