visual features
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
- Bridging the gap to real-world language-grounded visual concept learning
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic
- Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
- ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
- Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval