What Images Cannot Say: Language-Guided Olfactory Representation Learning
Abstract
Images tell us what a scene looks like, but rarely what itwould feel like to be there. While recent datasets pair visual sceneswith electronic-nose measurements, aligning smell signals with imagesremains challenging because many olfactory cues arise from contextualenvironmental factors that are not directly visible in pixels. We intro-duce SCENT, a multimodal framework that uses language guidance asa semantic bridge between vision and olfaction. Our approach leveragesVision-Language Models (VLMs) to generate scene descriptors captur-ing objects, environmental context, and plausible ambient smell cuessuggested by the visual scene. These descriptors provide semantic guid-ance for learning olfactory representations. We train a smell encoderthat maps electronic-nose signals into a shared embedding space alignedwith both visual and textual representations, and introduce a language-guided latent decomposition that separates object-specific odors fromcontextual environmental contributions. Experiments on the New YorkSmells dataset demonstrate that SCENT significantly improves cross-modal retrieval compared to vision-only baselines, achieving state-of-the-art performance on smell-to-image and smell-to-text retrieval tasks. Inaddition, our framework produces interpretable olfactory representationsthat enable the disentanglement of complex smell mixtures. Our resultsreveal the importance of contextual semantic information for groundingolfactory perception in multimodal learning and pave the way for futureresearch in this area.