Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Abstract
Predicting object dynamics (i.e., world modeling) is a fun-damental challenge for robotic manipulation, and modeling deformableobjects presents a particularly difficult case due to their high-dimensionalstate spaces and complex material properties. While current world mod-els approach this through two distinct paradigms: learning the dynamicsover the 2D pixel space or more explicit 3D geometric space. A systematicunderstanding of their relative strengths and limitations remains elusivedue to the lack of diverse, large-scale real-world data. To address this,we present Deform360, a large-scale visuotactile dataset featuring 198daily-life objects, 1,980 interaction sequences, and over 215 hours of ob-servations from 41 surround-view cameras and bimanual tactile grippersto capture both global motion and contact-induced local deformations.Leveraging a novel markerless visuotactile 3D tracking pipeline to ex-tract dense geometry and motion, we systematically evaluate currentstate-of-the-art world models, comparing 2D video models against 3Dparticle models. Finally, we provide a preliminary demonstration indi-cating the real-world applicability of our dataset by performing robotplanning tasks on deformable objects. Our analysis reveals key insightsinto the trade-offs between structural priors and scalability, providing asolid benchmark for future research in generalizable deformable object-centric world modeling. Project website: https://deform360.lhy.xyz