Scaffolding Dexterous Manipulation with Vision-Language Models

Gerhard Neumann (Karlsruhe Institute of Technology) · Vincent de Bakker (Karlsruher Institut für Technologie, Stanford University) · Joey Hejna (Stanford University) · Tyler Lum (Computer Science Department, Stanford University) · Onur Celik (ALR, Karlsruhe Institute of Technology KIT) · Aleksandar Taranovic (Karlsruhe Institute of Technology) · Denis Blessing (Karlsruher Institut für Technologie) · Jeannette Bohg (Stanford University) · Dorsa Sadigh (Stanford)
3d trajectory synthesisdexterous manipulationexploration guidancehigh-dimensional controlkeypoints identificationlow-level residual policyreal-world transferreference trajectoriesreinforcement learningsemantic knowledgesimulation experiencespatial knowledgetask-agnostic rewardstask-specific reward functionsvision-language models

Dexterous robotic hands are essential for performing complex manipulation tasks, yet remain difficult to train due to the challenges of demonstration collection and high-dimensional control. While reinforcement learning (RL) can alleviate the data bottleneck by generating experience in simulation, it typically relies on carefully designed, task-specific reward functions, which hinder scalability and generalization. Thus, contemporary works in dexterous manipulation have often bootstrapped from reference trajectories. These trajectories specify target hand poses that guide the exploration of RL policies and object poses that enable dense, task-agnostic rewards. However, sourcing suitable trajectories