inference efficiency
The speed and resource usage of making predictions with an AI model, critical in real-time applications and deployment scenarios.
- ARM: Adaptive Reasoning Model
- Activated LoRA: Fine-tuned LLMs for Intrinsics
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- CURE: Co-Evolving Coders and Unit Testers via Reinforcement Learning
- Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation
- Convex Potential Mirror Langevin Algorithm for Efficient Sampling of Energy-Based Models
- DISC: Dynamic Decomposition Improves LLM Inference Scaling
- DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- Diffusion on Demand: Selective Caching and Modulation for Efficient Generation
- Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
- Efficient Large Language Model Inference with Neural Block Linearization
- Efficient Rectified Flow for Image Fusion
- Efficient Speech Language Modeling via Energy Distance in Continuous Latent Space
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
- FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
- Generative Pre-trained Autoregressive Diffusion Transformer
- HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
- Linear Attention for Efficient Bidirectional Sequence Modeling
- Multi-head Temporal Latent Attention
- OSCAR: One-Step Diffusion Codec Across Multiple Bit-rates
- On the Hardness of Approximating Distributions with Tractable Probabilistic Models
- Parallel Scaling Law for Language Models
- PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
- Restoring Pruned Large Language Models via Lost Component Compensation
- SignFlow Bipartite Subgraph Network For Large-Scale Graph Link Sign Prediction
- SolidGeo: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
- SpectraLDS: Provable Distillation for Linear Dynamical Systems
- SymMaP: Improving Computational Efficiency in Linear Solvers through Symbolic Preconditioning
- Thinker: Learning to Think Fast and Slow
- Zebra-Llama: Towards Extremely Efficient Hybrid Models