WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild
Abstract
3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wildscenes, where drastic appearance variations and transient objects violatemulti-view consistency. Existing methods are fundamentally limited byindependent and discrete embeddings that struggle to capture continuousenvironmental changes or model spatially-varying local illumination. Toaddress these limitations, we propose WilLaGS, a unified frameworkfor robust 3D scene reconstruction and generative appearance synthe-sis under unconstrained settings. Specifically, we introduce a generativeappearance model where a β-VAE learns a structured and continuousmanifold of global appearance. Conditioned on the latent code, we con-struct a 3D neural appearance field that generates dynamic Tri-Plane fea-tures to encode spatially-varying local illumination effects. Furthermore,to suppress transient artifacts, we present a self-supervised perceptualmasking mechanism that leverages a Teacher-Student (EMA) architec-ture to derive a stable scene consensus, robustly identifying inconsistentregions via perceptual discrepancies. Extensive experiments on multipledatasets demonstrate that WilLaGS achieves state-of-the-art perfor-mance in reconstruction quality and novel view appearance synthesis,while maintaining real-time rendering efficiency.