Renrui Zhang
- AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
- Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
- Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow Reasoning
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- What We Miss Matters: Learning from the Overlooked in Point Cloud Transformers