inference latency
The time delay between submitting an input to an AI model and receiving the output. Inference latency is a critical consideration in real-time applications where immediate responses are required.
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- ASDSV: Multimodal Generation Made Efficient with Approximate Speculative Diffusion and Speculative Verification
- Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM Inference
- Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
- Diffusion on Demand: Selective Caching and Modulation for Efficient Generation
- FFN Fusion: Rethinking Sequential Computation in Large Language Models
- GSRF: Complex-Valued 3D Gaussian Splatting for Efficient Radio-Frequency Data Synthesis
- Inference-Time Hyper-Scaling with KV Cache Compression
- Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
- PhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth Observation
- R$^2$ec: Towards Large Recommender Models with Reasoning
- Scaling Embedding Layers in Language Models
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
- SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
- YOLOv12: Attention-Centric Real-Time Object Detectors
- msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyML