GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction
Abstract
Feed-forward 3D reconstruction provides an efficient paradigmfor scene modeling from image sequences. Scaling these models to largemonocular scenarios are constrained by excessive GPU memory foot-print, degraded local geometry, and long-term trajectory drift. Exist-ing chunk-based optimization strategies provide limited geometric con-straints and fail to maintain global consistency over extended trajecto-ries. We present a unified framework for stable and scalable feed-forward3D reconstruction from long monocular sequences. Our approach buildson coarse-to-fine trajectory alignment augmented by lightweight geomet-ric prior injection. Distilling monocular geometric cues into the feed-forward backbone via LoRA adaptation improves depth accuracy on finestructures while preserving inference efficiency. We introduce a hybrid-weight sparse ray-field optimization that leverages high-frequency geo-metric features to guide local point-cloud refinement and enforce con-sistent inter-frame ray constraints. Unlike prior chunk-based methods,this establishes strong cross-frame geometric coupling while maintainingscalability. Finally, an efficient trajectory stitching strategy with jointray-error optimization explicitly reduces accumulated drift. Extensiveexperiments show that our approach achieves competitive trajectory ac-curacy compared with representative SLAM systems, while maintainingglobally consistent 3D reconstruction in large-scale scenarios.