OmniX: Any-view and Any-time 4D reconstruction via Feed-forward Trajectory Fields
Abstract
Previous feed-forward 4D reconstruction methods either pre-dict per-frame static point clouds, ignoring foreground motion, or esti-mate point cloud trajectories while being limited to small camera mo-tions. This limits their ability to aggregate observations over time andreconstruct complete dynamic scenes under large viewpoint changes. Toaddress this, we propose OmniX, a feed-forward 4D reconstruction frame-work that predicts dense 3D point trajectories for every pixel fromvideos with large camera motion. 1) OmniX separates dynamic mo-tion modeling from static geometry prediction and represents motionwith a small set of dynamic tokens. Leveraging the sparse and low-rankstructure of 3D motion, these tokens generate trajectory fields for allpixels in all images while efficiently preserving global interactions. 2) Tofacilitate training, we build an automatic UE5-based 4D data engine andintroduce a dataset of 80k scenes and 1.28M multi-view videos with fullgeometric annotations. OmniX achieves state-of-the-art results on dense3D point trajectory prediction and 3D point tracking, with competitiveperformance on video depth estimation and camera pose estimation.