compression
The process of reducing the amount of data required to represent a dataset or model, often employed in AI to minimize storage needs and improve efficiency.
- A Partition Cover Approach to Tokenization
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- COALA: Numerically Stable and Efficient Framework for Context-Aware Low-Rank Approximation
- Explaining and Mitigating Crosslingual Tokenizer Inequities
- Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
- Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
- Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws