spatial perception
- 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
- Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
- From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
- GenSpace: Benchmarking Spatially-Aware Image Generation
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw