A Generalist Intracortical Motor Decoder

Leila Wehbe (Carnegie Mellon University) · Joel Ye (Carnegie Mellon University) · Fabio Rizzoglio (Northwestern University) · Xuan Ma (Northwestern University) · Adam Smoulder (CMU, Carnegie Mellon University) · Hongwei Mao (University of Pittsburgh) · Gary Blumenthal (University of Pittsburgh) · William Hockeimer (University of Pittsburgh) · Nicolas Kunigk (University of Pittsburgh) · Dalton Moore (University of Chicago) · Patrick Marino (Phantom Neuro) · Raeed Chowdhury · J. Patrick Mayo (University of Pittsburgh) · Aaron Batista (University of Pittsburgh) · Steven Chase · Michael Boninger (University of Pittsburgh) · Charles Greenspon (University of Chicago) · Andrew B Schwartz (University of Pittsburgh) · Nicholas Hatsopoulos (University of Chicago) · Lee Miller (Northwestern University at Chicago) · Kristofer Bouchard (Lawrence Berkeley National Laboratory) · Jennifer Collinger (University of Pittsburgh) · Robert Gaunt (University of Pittsburgh)
autoregressive transformercomplexity reductiondata integrationdownstream decoding tasksfoundation modelsintracortical microelectrode datamodel generalizationmotor covariatesmotor decodingneural datasetsneural distribution shiftsneural population spiking activityneurotechnologyoutput stereotypysensor variabilitysensorimotor neuroscience

Mapping the relationship between neural activity and motor behavior is a central aim of sensorimotor neuroscience and neurotechnology. While most progress to this end has relied on restricting complexity, the advent of foundation models instead proposes integrating a breadth of data as an alternate avenue for broadly advancing downstream modeling. We quantify this premise for motor decoding from intracortical microelectrode data, pretraining an autoregressive Transformer on 2000 hours of neural population spiking activity paired with diverse motor covariates from over 30 monkeys and humans. The resulting model is broadly useful, benefiting decoding on 8 downstream decoding tasks and generalizing to a variety of neural distribution shifts. However, we also highlight that scaling autoregressive Transformers seems unlikely to resolve limitations stemming from sensor variability and output stereotypy in neural datasets.