From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Abstract
Vision foundation models are capable of generalizing across3-dimensional (3D) scenes with high-fidelity estimates; their empiricalsuccess can be attributed to training on large-scale datasets of perspec-tive images. However, when transferred to wide field-of-view (FoV) im-ages, such as those captured by fisheye cameras, they return erroneousoutputs due to a covariate shift stemming from the radial distortion onthe image pixels. We propose a method to generalize vision foundationmodels to fisheye cameras. The crux of our method lies in a set of learn-able parameters, termed Distortion Extenders (DEX), that model thefisheye distortion coefficients and the distributional shift between fish-eye and perspective images encoded in the latent space. By minimizing aself-supervised alignment loss, DEX transforms the latent embeddings offisheye images to resemble those of perspective images to recover high-fidelity estimates. DEX is architecture- and task-agnostic: We demon-strate DEX on monocular depth estimation and open-vocabulary seg-mentation for convolution- and Transformer-based architectures, wherewe consistently improve over baselines across indoor and outdoor fisheyedatasets. As a byproduct, the activations of DEX can also be decoded todistortion coefficients to support camera calibration. Code available at:https://github.com/Suchisrit/DEX.