NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
multi-modal large language model
3 papers
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
Universal Visuo-Tactile Video Understanding for Embodied Interaction
VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model