multimodal benchmarks
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- Learning to Instruct for Visual Instruction Tuning
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions