structured pruning
Structured pruning goes beyond individual weight removal in neural networks and also entails the systematic removal of entire neurons, layers, or filters. This approach helps maintain the network’s architecture while enhancing efficiency.
- DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
- Elastic ViTs from Pretrained Models without Retraining
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
- Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- Restoring Pruned Large Language Models via Lost Component Compensation