textual descriptions
Verbal or written explanations that characterize objects, events, or concepts, often used as input for AI models to enhance understanding or generate relevant outputs.
- ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
- EchoShot: Multi-Shot Portrait Video Generation
- KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
- LLM-Driven Treatment Effect Estimation Under Inference Time Text Confounding
- Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
- Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
- TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence
- VLMs can Aggregate Scattered Training Patches