inference speedup
- BoltzNCE: Learning likelihoods for Boltzmann Generation with Stochastic Interpolants and Noise Contrastive Estimation
- Compress Large Language Models via Collaboration Between Learning and Matrix Approximation
- Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
- GAMMA: Gated Multi-hop Message Passing for Homophily-Agnostic Node Representation in GNNs
- StruDiCO: Structured Denoising Diffusion with Gradient-free Inference-stage Boosting for Memory and Time Efficient Combinatorial Optimization
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model