Bin Wang
- Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- LogicTree: Improving Complex Reasoning of LLMs via Instantiated Multi-step Synthetic Logical Data
- OmniTry: Virtual Try-On Anything without Masks
- ROSE: Remove Objects with Side Effects in Videos
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection