Lin Ma
- FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
- GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction-Editing Data and Long Captions
- VITRIX-UniViTAR: Unified Vision Transformer with Native Resolution