speculative decoding
An advanced decoding technique used in language generation to predict multiple potential sequences and then select the most likely one, improving the quality and coherence of generated text in models, particularly in large language models.
- ASDSV: Multimodal Generation Made Efficient with Approximate Speculative Diffusion and Speculative Verification
- Accelerating Diffusion LLMs via Adaptive Parallel Decoding
- AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
- Approximately Aligned Decoding
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
- EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization
- GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
- OmniDraft: A cross-vocabulary, online adaptive drafter for on-device speculative decoding
- STree: Speculative Tree Decoding for Hybrid State Space Models
- Scaling Speculative Decoding with Lookahead Reasoning
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- SpecEM: Training-Free LLM Ensembling via Iterative Drafting, Verification, and Online Feedback
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMs
- SpecMER: Fast Protein Generation with K-mer Guided Speculative Decoding
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
- Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
- SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
- TPP-SD: Accelerating Transformer Point Process Sampling with Speculative Decoding
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Traversal Verification for Speculative Tree Decoding
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
- Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding