multimodal llms
Large language models (LLMs) that can process and generate data across multiple modalities, such as text, images, and audio, improving their applicability to diverse tasks.
- Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
- Learning to Instruct for Visual Instruction Tuning
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Revealing Multimodal Causality with Large Language Models
- Synergistic Tensor and Pipeline Parallelism
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study