Jifeng Dai
- GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
- Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings