Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
Abstract
Recent advances in 3D feedforward reconstruction neural net-works have achieved remarkable success in dense reconstruction from imageswithout any camera parameters. Yet, equipping these models with robustsemantic understanding remains an open problem. Here we introduce anapproach that performs 3D reconstruction and 3D panoptic segmentationin a unified framework. We build on existing 3D reconstruction models andaugment them with a set-based mask decoder. The approach is jointly trainedwith a geometric and semantic loss, which are shown to be mutually beneficial.More precisely, the features are initialized from the geometric informationand then finetuned to capture jointly geometry and semantics. We demon-strate the generality of our approach by successfully applying our frameworkboth to online and all-to-all attention reconstruction backbones. Our methodachieves state-of-the-art performance in 3D panoptic segmentation acrossScanNet, ScanNet200, and ScanNet++ datasets. Ablation studies show thatsuch joint training of a unified model equips 3D feedforward reconstructionneural networks with panoptic segmentation and yields mutually beneficialimprovements.