PointSplat: Compact Gaussian Splatting via Human-Centric Prediction
Abstract
Producing 3D human representations from input views onthe fly is essential for immersive live streaming systems, where repre-sentation compactness is as critical as high fidelity given limited com-putational power and transmission bandwidth. Although recent feed-forward reconstruction methods achieve impressive quality through theview-centric prediction of 3D representations, they repeatedly encode thesame subject content across multiple views, leading to significant inter-view redundancy. Our key insight is to perform predictions directly in 3Dspace, enabling the network to learn and produce a highly compact rep-resentation. To this end, we propose PointSplat, a novel human-centricapproach that directly infers Gaussian primitives from an input pointset. The proposed method first estimates a coarse geometric proxy andperforms ray casting to prune redundant points and establish explicit2D–3D correspondences. Subsequently, it employs a Point-Image Trans-former to fuse appearance and geometry features, predicting Gaussianattributes in a single forward pass. This design restricts predictions toforeground regions of interest, substantially reducing the total numberof Gaussians while improving novel-view rendering quality. Extensiveexperiments demonstrate that PointSplat achieves higher efficiency andquality while exhibiting strong robustness to variations in view countand image resolution across multiple datasets.