Unsupervised Point Cloud Registration via Training-Time Semantic Guidance
Abstract
Unsupervised registration of large-scale LiDAR point cloudsremains challenging due to the geometric ambiguity inherent in out-door scenes, which degrades pseudo-label quality and leads to subop-timal convergence, particularly for sparse, low-resolution scans such asthose from nuScenes. We reveal that registration models intrinsicallyencode semantic awareness that strongly correlates with registration ac-curacy, albeit without explicit semantic supervision. However, this nativeawareness is fragile: noisy supervision arising from geometric ambigu-ity in unsupervised settings rapidly erodes the learned semantic struc-ture, causing performance collapse. To this end, we propose CAESAR,a teacher-student framework guided by an off-the-shelf 3D segmenta-tion model exclusively during training. We observe that potential in-lier matches are often buried just beneath a few spurious neighbors inthe noisy feature space, motivating Dual-Cue Guided Re-Matching torecover them through reselection rather than simply rejecting. Build-ing on this, a train-only Semantic-Geometric Label Mining performsCorresponding author.lightweight, batch-specific teacher refinement and mines reliable pseudo-labels under semantic guidance. We further introduce Semantic Predic-tive Distillation to consolidate the student’s semantic awareness in thefeature space. Extensive experiments on KITTI and nuScenes demon-strate state-of-the-art performance, with pronounced gains on the chal-lenging nuScenes benchmark. Crucially, CAESAR incurs zero inferenceoverhead and requires no semantic annotations on the registra-tion data.