FlowPrune: Accelerating Attention Flow Calculation by Pruning Flow Network
attention flowattention graphattention mechanismscomputational complexityedge pruningflowpruneinfluential input tokensinterpretability techniqueslarge multimodal modelslayer compressionmax-flow algorithmsmax-flow min-cut theoremrobustness evaluationtransformer architecture
The Transformer architecture serves as the foundation of modern AI systems, powering recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs). Central to these models, attention mechanisms capture contextual dependencies via token interactions. Beyond inference, attention has been widely adopted for interpretability, offering insights into model behavior. Among interpretability techniques, attention flow