World Models for Learning Dexterous Hand-Object Interactions from Human Videos
Abstract
Modeling dexterous hand–object interactions is challengingas it requires understanding how subtle finger motions influence theenvironment through contact with objects. While recent world mod-els address interaction modeling, they typically rely on coarse actionspaces that fail to capture fine-grained dexterity. We, therefore, intro-duce DexWM, a Dexterous Interaction World Model that predicts futurelatent states of the environment conditioned on past states and dexterousactions. To overcome the scarcity of finely annotated dexterous datasets,DexWM represents actions using finger keypoints extracted from ego-centric videos, enabling training on over 900 hours of human and non-dexterous robot data. Further, to accurately model dexterity, we find thatpredicting visual features alone is insufficient; therefore, we incorporatean auxiliary hand consistency loss that enforces accurate hand config-urations. DexWM outperforms prior world models conditioned on text,navigation, or full-body actions in future-state prediction and demon-strates strong zero-shot transfer to unseen skills on a Franka Panda armwith an Allegro gripper, surpassing Diffusion Policy by over 50% on av-erage across grasping, placing, and reaching tasks. Codes and dataset areavailable on the project page: https://raktimgg.github.io/dexwm/.