memory consumption
This term refers to the amount of memory resources utilized by a model during execution, which can impact performance and scalability. Efficient memory consumption is crucial for deploying AI models, especially on resource-constrained devices.
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression
- CALM-PDE: Continuous and Adaptive Convolutions for Latent Space Modeling of Time-dependent PDEs
- LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- PolarQuant: Leveraging Polar Transformation for Key Cache Quantization and Decoding Acceleration
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs