inference costs
The computational and resource expenses associated with running a trained model to make predictions or generate outputs. It encompasses processing time, memory usage, and energy consumption, and is critical in evaluating the practical deployment of AI systems.
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- DecompNet: Enhancing Time Series Forecasting Models with Implicit Decomposition
- Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
- Holistic Order Prediction in Natural Scenes
- How Far Are We from Optimal Reasoning Efficiency?
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- Know What You Don't Know: Uncertainty Calibration of Process Reward Models
- Language Models can Self-Improve at State-Value Estimation for Better Search
- Learned Prefix Caching for Efficient LLM Inference
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- Neural Attention Search
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
- SymMaP: Improving Computational Efficiency in Linear Solvers through Symbolic Preconditioning
- Theoretical Benefit and Limitation of Diffusion Language Model
- Training Language Models to Reason Efficiently
- VeriThinker: Learning to Verify Makes Reasoning Model Efficient