Rosetum3D: A Large-Scale 3D Vision Dataset from Preharvest Roses
Abstract
The global rose cultivation industry has experienced continued expansion in recent years. There has been a set of value evaluation criteria covering multiple dimensions for roses. However, the evaluations are mainly conducted by human experts after harvesting, which is time-consuming and labor-intensive. 3D computer vision technologies are promising to automate the process before harvesting, but the lack of large-scale datasets with localization annotations that capture plant architecture hinders the progress. To bridge this gap, we propose Rosetum3D, which is a large-scale 3D vision dataset for preharvest roses. The dataset is constructed via occlusion-robust multi-view RGB-D capture protocols in commercial greenhouses. We provide fine-grained 2D localization labels using bounding boxes and botanically defined keypoints, and then obtain the 3D structures recovered through depth backprojection. Rosetum3D contains 21,114 images and 46,848 annotated rose objects. Models trained on Rosetum3D have achieved 2D/3D rose localization, which is a crucial step for automated preharvest quality grading and growth monitoring. Beyond localization, Rosetum3D serves as a benchmark for agricultural vision tasks, including 2D rose object detection, local feature matching, depth estimation, and instance reidentification. By enabling data-driven precision agriculture, Rosetum3D paves the way for robotic harvesting systems and AI-driven yield prediction in protected cultivation. The dataset are available at https: //huggingface.co/datasets/WaterMelon2333/Rosetum3D/tree/main.