Bipolar Self-attention for Spiking Transformers

Dehao Zhang (University of Electronic Science and Technology of China) · Malu Zhang (National University of Singapore) · Shuai Wang (Nanjing University) · Jingya Wang (ShanghaiTech University) · Zeyu Ma (University of Electronic Science and Technology of China) · Yang Yang (Nanjing University of Science and Technology) · Haizhou Li (The Chinese University of Hong Kong (Shenzhen); National University of Singapore) · Jieyuan (Eric) Zhang (University of Electronic Science and Technology of China) · Yimeng Shan (Liaoning Technical University) · Yichen Xiao (University of Electronic Science and Technology of China) · Honglin Cao (University of Electronic Science and Technology of China) · Haonan Zhang (University of Electronic Science and Technology of China)
attention allocationbipolar self-attentionenergy-efficient architecturesheteropolar interactionshomopolar interactionsimage classificationlow-entropy activationmulti-polar membrane potentialreal-valued computationrow-stochasticitysemantic segmentationshiftmax approximationsoftmax functionsspiking neural networksspiking self-attentionternary matrix multiplication

Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through comprehensive analysis, we attribute this gap to these two factors. First, the binary nature of spike trains limits Spiking Self-attention (SSA)’s capacity to capture negative–negative and positive–negative membrane potential interactions on Querys and Keys. Second, SSA typically omits Softmax functions to avoid energy-intensive multiply-accumulate operations, thereby failing to maintain row-stochasticity constraints on attention scores. To address these issues, we propose a Bipolar Self-attention (BSA) paradigm, effectively modeling multi-polar membrane potential interactions with a fully spike-driven characteristic. Specifically, we demonstrate that ternary matrix multiplication provides a closer approximation to real-valued computation on both distribution and local correlation, enabling clear differentiation between homopolar and heteropolar interactions. Moreover, we propose a shift-based Softmax approximation named Shiftmax, which efficiently achieves low-entropy activation and partly maintains row-stochasticity without non-linear operation, enabling precise attention allocation. Extensive experiments show that BSA achieves substantial performance improvements across various tasks, including image classification, semantic segmentation, and event-based tracking. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers.