SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
Abstract
VAR has gained widespread popularity due to its next-scaleprediction paradigm. However, it faces substantial performance bottle-necks when handling complex scenes with multiple objects and attributes.Existing diffusion-based enhancement methods fail to adequately addressthe unique challenge of cross-scale error propagation and accumulationin VAR. To this end, we propose SynVAR, the first training-free enhance-ment framework specifically tailored for the VAR paradigm, which intro-duces a spatial-semantic collaborative control strategy to effectively sup-press propagation error and improve generation quality. SynVAR com-prises three key components: (1) Global guidance to ensure reasonablespatial structure in the early stages, (2) Receptive field constraints tomitigate early-stage semantic confusion, (3) High-frequency compensa-tion to recover fine-grained details. Extensive quantitative and qualita-tive experiments demonstrate the significant improvements in the abilityof SynVAR to enhance the VAR’s capability for complex scene modeling.