DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models
Abstract
Despite remarkable progress in text-guided image editing, gen-erative models frequently fail to preserve visual object consistency, definedas the preservation of a subject’s key attributes throughout the editingprocess. We address this limitation through three contributions. First, weintroduce ABO-Edit, a dataset specifically designed to study object con-sistency, comprising over 12,000 triplets of source images, editing prompts,and high-quality target images rendered from artist-designed 3D assets,with multi-view coverage and human-verified quality control. Second, weuncover an overlooked property of image-editing rectified flow models: theconditioning embedding space, not directly supervised during training,encodes a prediction of the final generated image even at high noise levels.Third, exploiting this finding, we propose FlowMirror, a parameter-freeauxiliary loss that supervises this conditioning embedding space. Withoutarchitectural changes, our method improves generation quality acrossseveral metrics over baselines. Page: francescotaioli.github.io/DiTailed