Jiahui Zhang
- 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
- Dynamic and Chemical Constraints to Enhance the Molecular Masked Graph Autoencoders
- Enhancing the Maximum Effective Window for Long-Term Time Series Forecasting
- From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D