model compression
Techniques aimed at reducing the size and computational demand of AI models while maintaining their performance. This is important for deploying models in resource-constrained environments or improving inference speed.
- $\text{S}^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
- 3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMs
- Activation-Informed Merging of Large Language Models
- Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
- Attribution-Driven Adaptive Token Pruning for Transformers
- Dynamical Low-Rank Compression of Neural Networks with Robustness under Adversarial Attacks
- Global Minimizers of $\ell^p$-Regularized Objectives Yield the Sparsest ReLU Neural Networks
- Hankel Singular Value Regularization for Highly Compressible State Space Models
- Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language Models
- MODEL SHAPLEY: Find Your Ideal Parameter Player via One Gradient Backpropagation
- PLD: A Choice-Theoretic List-Wise Knowledge Distillation
- Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
- Unified Scaling Laws for Compressed Representations