Stitch and Tell: A Structured Data Augmentation Method for Spatial Understanding

Peiwen Yuan (Beijing Institute of Technology) · Yiwei Li (Beijing Institute of Technology) · Shaoxiong Feng (Beijing Institute of Technology) · Jiayi Shi (Beijing Institute of Technology) · Prof. Kan · Yin Hang · Xiaomin He (Peking University) · Wenxiao Fan (Beijing Institute of Technology)
annotation-free methodasymmetric propertiesgeneral vision-language benchmarkshalvallavamultimodal dataperformance improvementquestion answer pairsspatial hallucinationsspatial understanding tasksspatially-aware captionsspatially-aware structurestitched image-text pairsstructured spatial supervisiontraining datasetsvision-language models

Existing vision-language models often suffer from spatial hallucinations, i.e., generating incorrect descriptions about the relative positions of objects in an image. We argue that this problem mainly stems from the asymmetric properties between images and text. To enrich the spatial understanding ability of vision-language models, we propose a simple, annotation-free, plug-and-play method named Stitch and Tell (abbreviated as SiTe), which injects structured spatial supervision into multimodal data. It constructs stitched image–text pairs by stitching images along a spatial axis and generating spatially-aware captions or question answer pairs based on the layout of stitched image, without relying on costly advanced models or human involvement. We evaluate SiTe across three architectures including LLaVA-v1.5-7B, LLaVA-Qwen2-1.5B and HALVA-7B, two training datasets, and thirteen benchmarks. Experiments show that SiTe improves spatial understanding tasks such as $\text{MME}_{\text{Position}}$ (+5.50\%) and Spatial-MM (+4.19\%), while maintaining or improving performance on general vision-language benchmarks. Our findings suggest that explicitly injecting spatially-aware structure into training data offers an effective way to mitigate spatial hallucinations and improve spatial understanding, while preserving general vision-language capabilities.