See & Sniff: Learning Visuo-Olfactory Representations
Abstract
While modern multimodal models integrate vision with lan-guage, audio, or touch, olfaction remains largely unexplored due to thelack of paired visuo-olfactory data. We introduce SmellNet-V, a scal-able visuo-olfactory dataset built on the insight that odor identity islargely invariant to visual transformations within a semantic category.This allows us to synthetically pair smell-only samples with semanticallyaligned in-the-wild web images, converting a unimodal olfactory datasetinto a cross-modal benchmark without costly co-collection. Building onthis dataset, we propose See & Sniff, a self-supervised framework thatlearns joint visuo–olfactory representations via dense local alignment andnaturally produces smell saliency maps for spatial grounding of odorsources. We further introduce pixel-level smell localization task and abenchmark for evaluation. Our method surpasses smell-only baselines by7% in smell classification from smell alone and generalizes to cross-modalretrieval and smell localization, establishing visuo-olfactory learning asa new direction in multimodal perception.