Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

Dehao Zhang (University of Electronic Science and Technology of China) · Malu Zhang (National University of Singapore) · Shuai Wang (Nanjing University) · Jingya Wang (ShanghaiTech University) · Wenjie Wei (University of Electronic Science and Technology of China) · Zeyu Ma (University of Electronic Science and Technology of China) · Guoqing Wang (University of Electronic Science and Technology of China) · Yang Yang (Nanjing University of Science and Technology) · Haizhou Li (The Chinese University of Hong Kong (Shenzhen); National University of Singapore)
adaptive threshold mechanismcomputational efficiencydendritic resonate-and-fire modeldendritic structureedge platformseffective memory capacityenergy efficiencyfrequency representationhistorical spiking activitylong sequence modelingmulti-dendritic architectureresonate-and-fire neuronssparse spikesspatiotemporal spike trainstraining speed

The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiotemporal spike trains, making them well-suited for long sequence modeling. However, RF neurons exhibit limited effective memory capacity and a trade-off between energy efficiency and training speed on complex temporal tasks. Inspired by the dendritic structure of biological neurons, we propose a Dendritic Resonate-and-Fire (D-RF) model, which explicitly incorporates a multi-dendritic and soma architecture. Each dendritic branch encodes specific frequency bands by utilizing the intrinsic oscillatory dynamics of RF neurons, thereby collectively achieving comprehensive frequency representation. Furthermore, we introduce an adaptive threshold mechanism into the soma structure. his mechanism adjusts the firing threshold according to historical spiking activity, thereby reducing redundant spikes while maintaining training efficiency in long-sequence tasks. Extensive experiments demonstrate that our method maintains competitive accuracy while substantially ensuring sparse spikes without compromising computational efficiency during training. These results underscore its potential as an effective and efficient solution for long sequence modeling on edge platforms.