CulinaryCut: A Physics-aware Vision-Language-Action Benchmark for Food Cutting via Material Point Method
Abstract
Food cutting is a representative non-rigid manipulation taskinvolving deformation, contact forces, and material separation, yet itRobot Simulation remains largely unexplored in VLA research. Existing robot manipula-tion datasets are primarily built around rigid objects and struggle tojointly capture the spatial precision, deformation, topology changes, andCut the banana force interactions required for cutting. To systematically study theseat the center. challenges, we introduce CulinaryCut, a benchmark that integrates anniskill MPM-based deformable simulator with a robot environment to generatedataset rather than provided as policy inputs, enabling evaluation of ex-isting VLAs while exposing where physical grounding is needed. Usingthree representative VLA baselines, we identify two core limitations innon-rigid cutting. First, a geometry gap: VLAs struggle to map ratio-and direction-based instructions onto object geometry, especially whensequential cuts change the object’s topology. Second, a physics gap: be-cause policies output motion without a notion of material stiffness orcontact resistance, a geometrically accurate path may still fail to severthe object. We show that training VLAs on physics-grounded trajecto-ries, without modifying the policy architecture or adding force inputs,improves cutting success and sim-to-real transfer. Together, CulinaryCutprovides a foundation for evaluating spatial reasoning and physical plau-sibility in deformable manipulation.