FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
Abstract
We present FLEG, a feed-forward network that reconstructslanguage-embedded 3D Gaussians from multi-view images. Previous feed-forward language-embedded Gaussian reconstruction methods are re-stricted to a fixed number of input views and typically attach a language-aligned semantic embedding to each Gaussian, resulting in impracticalinput settings and semantic redundancy. In contrast, we introduce ageometric-semantic dual-branch distillation framework that supports avariable number of uncalibrated and unposed multi-view images as input.We also propose a novel-view-based distillation strategy during trainingthat mitigates overfitting to input views. In addition, we observe that se-mantic representations are significantly sparser than geometric ones, andper-Gaussian language embedding is unnecessary. To exploit this spar-sity, we design a decoupled language embedding strategy that representslanguage information with a sparse set of semantic Gaussians, ratherthan attaching embeddings to every Gaussian. Compared with densepixel-aligned per-Gaussian embedding schemes, our method uses only5% of the language embeddings while maintaining comparable semanticfidelity, effectively reducing storage costs. Extensive experiments demon-strate that FLEG outperforms state-of-the-art feed-forward reconstruc-tion and language-embedded Gaussian methods in both reconstructionquality and language-aligned semantic representation.