Robust onion: Peeling Open Vocab Object Detectors Under Noise
Abstract
The impact of real-world noise on Open Vocabulary ObjectDetectors (OV-ODs) remains poorly understood due to their architec-tural complexity. We present our comprehensive analysis, Robust Onion,an empirical study that uses controlled synthetic visual degradations topeel OV-ODs layer-by-layer, revealing how, why, and where robustnessdegrades, systematically analyzing feature collapse. Our findings revealthat models with similar vision backbones exhibit comparable robust-ness, driven by similar feature collapse at similar layers, while factorssuch as pretraining strategy, architectural nuances, and caption supervi-sion contribute little. Robustness is primarily governed by the image do-main rather than annotations, explaining the similar robustness impacton COCO and LVIS, and why datasets like ODinW-13 can give an im-pression of inflated robustness due to large, isolated objects. Finally, wevalidate our insights by improving robustness on real-world BDD-100K,WiderFace, and VisDRONE via our lightweight plug-and-play NN &TK0 approach, using 96× fewer trainable parameters than end-to-endtraining. We also explain the prior works’ robustness observations.