Intend to Move: A Multimodal Dataset for Intention-Aware Human Motion Understanding

Ryo Umagami (The University of Tokyo) · Liu Yue (The University of Tokyo, Tokyo University) · Xuangeng Chu (The University of Tokyo) · Ryuto Fukushima (The University of Tokyo, The University of Tokyo) · Tetsuya Narita (The University of Tokyo) · Yusuke Mukuta (The University of Tokyo) · Tomoyuki Takahata (Tokyo Denki University) · Jianfei Yang (Nanyang Technological University) · Tatsuya Harada (The University of Tokyo / RIKEN)
3d motion sequencescausal factorsembodied intelligenceembodied reasoninghigh-level goalsintention-groundedintentionalitykinematicslanguage annotationsmotion modelingmultimodal datasetphysically coherent motionrgb-d videoscene geometrysemantic factorssocially coherent motion

Human motion is inherently intentional, yet most motion modeling paradigms focus on low-level kinematics, overlooking the semantic and causal factors that drive behavior. Existing datasets further limit progress: they capture short, decontextualized actions in static scenes, providing little grounding for embodied reasoning. To address these limitations, we introduce $\textit{Intend to Move (I2M)}$, a large-scale, multimodal dataset for intention-grounded motion modeling. I2M contains 10.1 hours of two-person 3D motion sequences recorded in dynamic realistic home environments, accompanied by multi-view RGB-D video, 3D scene geometry, and language annotations of each participant’s evolving intentions. Benchmark experiments reveal a fundamental gap in current motion models: they fail to translate high-level goals into physically and socially coherent motion. I2M thus serves not only as a dataset but as a benchmark for embodied intelligence, enabling research on models that can reason about, predict, and act upon the ``why'' behind human motion.