Hengshuang Zhao
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- GenSpace: Benchmarking Spatially-Aware Image Generation
- LiteReality: Graphic-Ready 3D Scene Reconstruction from RGB-D Scans
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- PlayerOne: Egocentric World Simulator
- PlayerOne: Egocentric World Simulator
- ROSE: Remove Objects with Side Effects in Videos
- Seg-VAR:Image Segmentation with Visual Autoregressive Modeling
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance