FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network

Xin Geng (Southeast University) · Xu Yang (Microsoft) · Yu Chen (Shanghai Jiaotong University) · Shuo Xu (University of Maryland, College Park) · Shuxia Lin (Southeast University)
attention flowattention graphattention mechanismscomputational complexityedge pruningflowpruneinfluential input tokensinterpretability techniqueslarge multimodal modelslayer compressionmax-flow algorithmsmax-flow min-cut theoremrobustness evaluationtransformer architecture

The Transformer architecture serves as the foundation of modern AI systems, powering recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs). Central to these models, attention mechanisms capture contextual dependencies via token interactions. Beyond inference, attention has been widely adopted for interpretability, offering insights into model behavior. Among interpretability techniques, attention flow