Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
Abstract
Recent advances in Virtual Try-On (VTON) and VirtualTry-Off (VTOFF) have greatly improved photo-realistic fashion synthesisand garment reconstruction. However, existing datasets remain static,lacking instruction-driven editing for controllable and interactive fash-ion generation. In this work, we introduce the Dress Editing Dataset(Dress-ED), the first large-scale benchmark that unifies VTON, VTOFF,and text-guided garment editing within a single framework. Each sam-ple in Dress-ED includes an in-shop garment image, the correspondingperson image wearing the garment, their edited counterparts, and anatural-language instruction of the desired modification. Built througha fully automated multimodal pipeline that integrates MLLM-basedgarment understanding, diffusion-based editing, and LLM-guided ver-ification, Dress-ED comprises over 146k verified quadruplets spanningthree garment categories and seven edit types, including both appear-ance (e.g., color, pattern, material) and structural (e.g., sleeve length,neckline) modifications. Based on this benchmark, we further propose aunified multimodal diffusion framework that jointly reasons over linguisticinstructions and visual garment cues, serving as a strong baseline forinstruction-driven VTON and VTOFF.