SignRefine: Adapting Foundational Video Models for Sign Language Generation
Abstract
Sign language video generation demands precise hand andfacial articulation, yet modern video diffusion models, trained predom-inantly on spoken-language video, produce artifacts that render signingunintelligible. We propose SignRefine, a sign language video genera-tion model that produces comprehensible signing from 2D keypoint con-ditioning alone, generalizing across appearances and visual conditions.Our approach builds on a pretrained video diffusion transformer and in-troduces local adapters with spatial grounding to selectively refine handand face regions, steering the strong base model’s prior toward accu-rate articulation. To enable this work and support broader sign languageresearch, we present NVSign, a large-scale dataset of video content na-tively produced in sign language, offering diverse signer appearances,environments, and natural conversational settings. Trained on this data,our model shows up to 30% improvement in hand pose precision metricsover the strongest baseline and is preferred by sign language users forvisual quality and comprehensibility in more than 80% of comparisons.