Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning
Abstract
Prevailing 2D-centric visuomotor policies exhibit a pronounceddeficiency in novel view generalization, as their reliance on static observa-tions hinders consistent action mapping across unseen views. In response,we introduce GenSplat, a feed-forward 3D Gaussian Splatting frameworkthat facilitates view-generalized policy learning through novel view ren-dering. GenSplat employs a permutation-equivariant architecture to re-construct high-fidelity 3D scenes from sparse, uncalibrated inputs in asingle forward pass. To ensure structural integrity, we design a 3D-priordistillation strategy that regularizes the 3DGS optimization, preventingthe geometric collapse typical of purely photometric supervision. By ren-dering diverse synthetic views from these stable 3D representations, wesystematically augment the observational manifold during training. Thisaugmentation forces the policy to ground its decisions in underlying 3Dstructures, thereby ensuring robust execution under severe spatial per-turbations where baselines severely degrade.