modality gap
The differences and challenges that arise when models trained on one type of modality (like text) are applied to another (like images), often complicating transfer learning.
- Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation
- GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap Data
- Global Minimizers of Sigmoid Contrastive Loss
- Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
- HMVLM:Human Motion-Vision-Language Model via MoE LoRA
- Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
- ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
- Where Does It Exist from the Low-Altitude: Spatial Aerial Video Grounding