NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
latency overhead
The additional delay incurred during processing or response times in AI systems, which can impact user experience and system performance.
3 papers
FlashMoE: Fast Distributed MoE in a Single Kernel
Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization