NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
memory reduction
3 papers
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
PseuZO: Pseudo-Zeroth-Order Algorithm for Training Deep Neural Networks