NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
token distribution
3 papers
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
Foundations of Top-$k$ Decoding for Language Models
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache