multimodal fusion
The integration of information from multiple modalities (e.g., text, image, audio) to improve the overall performance of a model, leveraging different types of data to make richer predictions.
- Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet Process
- CogPhys: Assessing Cognitive Load via Multimodal Remote and Contact-based Physiological Sensing
- CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- MolVision: Molecular Property Prediction with Vision Language Models
- Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning
- VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations