multi-modal learning
- CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems
- COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
- DAVE: Diagnostic benchmark for Audio Visual Evaluation
- Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation
- OctoNet: A Large-Scale Multi-Modal Dataset for Human Activity Understanding Grounded in Motion-Captured 3D Pose Labels