Toward Human Deictic Gesture Target Estimation

James Rehg (University of Illinois at Urbana-Champaign) · Xu Cao (University of Illinois at Urbana-Champaign & PediaMed AI) · Pranav Virupaksha (Georgia Institute of Technology) · Sangmin Lee (Korea University) · Bolin Lai (Georgia Tech) · Wenqi Jia (University of Illinois at Urbana-Champaign) · Jintai Chen
co-speech deictic gesturescomputer visionemerging research areagaze-aware joint cross attentiongaze-following cuesgesture existence predictiongesture semantic target estimationgesture target estimationgesture target mask predictiongesture target predictionhuman communicationintention inferencemultimodal integrationsocial interactionspatial location

Humans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central role in human communication. Understanding the intended targets of another individual's deictic gestures enables inference of their intentions, comprehension of their current actions, and prediction of upcoming behaviors. Despite its significance, gesture target estimation remains an underexplored task within the computer vision community. In this paper, we introduce GestureTarget, a novel task designed specifically for comprehensive evaluation of social deictic gesture semantic target estimation. To address this task, we propose TransGesture, a set of Transformer-based gesture target prediction models. Given an input image and the spatial location of a person, our models predict the intended target of their gesture within the scene. Critically, our gaze-aware joint cross attention fusion model demonstrates how incorporating gaze-following cues significantly improves gesture target mask prediction IoU by 6% and gesture existence prediction accuracy by 10%. Our results underscore the complexity and importance of integrating gaze cues into deictic gesture intention understanding, advocating for increased research attention to this emerging area. All data, code will be made publicly available upon acceptance. Code of TransGesture is available at GitHub.com/IrohXu/TransGesture.