multimodal large language models
A class of models, such as those that integrate text and visual information, designed to understand and generate content across various modalities, enhancing their versatility in tasks.
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
- AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
- Adaptive Gradient Masking for Balancing ID and MLLM-based Representations in Recommendation
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
- AnomalyCoT: A Multi-Scenario Chain-of-Thought Dataset for Multimodal Large Language Models
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- Backdoor Cleaning without External Guidance in MLLM Fine-tuning
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention Reallocation
- Boosting Knowledge Utilization in Multimodal Large Language Models via Adaptive Logits Fusion and Attention Reallocation
- CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
- CHASM: Unveiling Covert Advertisements on Chinese Social Media
- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization
- Counterfactual Evolution of Multimodal Datasets via Visual Programming
- Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models
- Doctor Approved: Generating Medically Accurate Skin Disease Images through AI-Expert Feedback
- Don't Just Chase “Highlighted Tokens” in MLLMs: Revisiting Visual Holistic Context Retention
- DreamPRM: Domain-reweighted Process Reward Model for Multimodal Reasoning
- DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
- EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
- EgoBlind: Towards Egocentric Visual Assistance for the Blind
- EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
- Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language Models
- FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
- ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection
- GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
- Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
- Improving Task-Specific Multimodal Sentiment Analysis with General MLLMs via Prompting
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
- InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
- Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
- Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- Lie Detector: Unified Backdoor Detection via Cross-Examination Framework
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
- MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation
- MLLM-ISU: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models based Intrusion Scene Understanding
- MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
- MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
- MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
- MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- MedSG-Bench: A Benchmark for Medical Image Sequences Grounding
- Meta CLIP 2: A Worldwide Scaling Recipe
- MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
- MobileUse: A Hierarchical Reflection-Driven GUI Agent for Autonomous Mobile Operation
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- Multimodal Tabular Reasoning with Privileged Structured Information
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- NavBench: Probing Multimodal Large Language Models for Embodied Navigation
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
- OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
- Omni-Mol: Multitask Molecular Model for Any-to-any Modalities
- OmniBench: Towards The Future of Universal Omni-Language Models
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
- Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level Guidance
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
- Scaling Language-centric Omnimodal Representation Learning
- Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
- Seeking and Updating with Live Visual Knowledge
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- SolidGeo: Measuring Multimodal Spatial Math Reasoning in Solid Geometry
- SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMs
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- Spatially-aware Weights Tokenization for NeRF-Language Models
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
- Structure-Aware Cooperative Ensemble Evolutionary Optimization on Combinatorial Problems with Multimodal Large Language Models
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
- The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
- UMU-Bench: Closing the Modality Gap in Multimodal Unlearning Evaluation
- UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
- Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- Unlocking Multimodal Mathematical Reasoning via Process Reward Model
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction-Editing Data and Long Captions
- VLForgery Face Triad: Detection, Localization and Attribution via Multimodal Large Language Models
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought
- VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task Awareness
- Vid-SME: Membership Inference Attacks against Large Video Understanding Models
- Video-R1: Reinforcing Video Reasoning in MLLMs
- VideoCAD: A Dataset and Model for Learning Long‑Horizon 3D CAD UI Interactions from Video
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- Vision Function Layer in Multimodal LLMs
- Visual Instruction Bottleneck Tuning
- VisualLens: Personalization through Task-Agnostic Visual History
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLM