inference optimization
Inference optimization involves enhancing the efficiency and speed of model predictions. Techniques may include quantization, pruning, or using specialized hardware to reduce the computational load during the inference phase.
- Fast Inference for Augmented Large Language Models
- KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- Multilevel neural simulation-based inference
- R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
- Sim-LLM: Optimizing LLM Inference at the Edge through Inter-Task KV Reuse