visual tasks
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
- Towards Interpretable and Efficient Attention: Compressing All by Contracting a Few
- Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
- Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations