inference acceleration
Techniques and strategies aimed at reducing the time and computational costs associated with making predictions with AI models. This is essential for deploying models in real-time applications.
- $\text{S}^2$Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
- DuSA: Fast and Accurate Dual-Stage Sparse Attention Mechanism Accelerating Both Training and Inference
- EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization
- Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
- Language Models Can Predict Their Own Behavior
- MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
- Plug-and-Play Context Feature Reuse for Efficient Masked Generation
- Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
- Radial Attention: $\mathcal O(n \log n)$ Sparse Attention for Long Video Generation
- Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
- SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
- Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
- TV-Rec: Time-Variant Convolutional Filter for Sequential Recommendation
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
- dKV-Cache: The Cache for Diffusion Language Models