NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
inference throughput
3 papers
Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs