UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment
Abstract
In this paper, we address the challenging problem of 4D re-construction from sparse-view videos. This setup usually relies on monoc-ular depth estimation to provide priors for the reconstruction model. Akey challenge arises from limited cross-view overlap and temporal vari-ation, making monocular depth predictions inconsistent across viewsand time. Existing methods align spatial and temporal dimensions inseparate stages, requiring foreground segmentation masks while failingto leverage temporal cues for cross-view alignment. Contrary to thesemethods, we propose a unified spatial-temporal depth alignment frame-work that jointly resolves cross-view and cross-time inconsistencies with-out distinguishing foreground/background. Our method represents depthmaps across views and time as a set of spatio-temporal neural fields.This representation not only yields fast convergence, but also capturesspatio-temporal correlation among depth maps implicitly, without de-pendence on external segmentation/tracking models. We also propose amulti-view depth-order loss while leveraging the classic scale-and-shift-invariant loss to further improve the final depth quality. The aligneddepths initialize and supervise Gaussian splatting models for 4D recon-struction. Experiments on Ego-Exo4D and EgoHuman demonstrate thatour improved depth alignment substantially benefits dynamic Gaussian-splatting-based reconstruction methods for novel-time/view synthesis andgeometry accuracy/consistency.