DASAM3D: A Unified Foundation Model for Enhanced 3D Scene Reconstruction and Segmentation
Abstract
We present DASAM3D, a framework that fuses the complementary strengths of SAM3D and Depth Anything 3 (DA3) for objectaware 3D scene reconstruction. A fundamental appearance–geometry trade-off exists: SAM3D yields geometrically coherent per-object Gaussian primitives but with unrealistic appearance, while DA3 delivers photorealistic 3DGS reconstructions yet lacks object-level geometric reasoning. Our key insight is to use SAM3D exclusively as a 3D geometric layout provider, guiding DA3 geometry refinement while preserving its photorealistic textures. DASAM3D proceeds through three differentiable stages—(1) Global Alignment Optimization, (2) Per-Object Alignment Refinement, and (3) Layout-Based Geometry Optimization—and inherits DA3’s flexibility for posed or pose-free inputs at any scale. Experiments on DL3DV-10K and MipNeRF-360 show consistent gains in reconstruction accuracy and novel view synthesis; a user-preference study confirms DASAM3D’s object renderings are preferred over SAM3D’s in visual fidelity.