Generative Relightable Avatars
Abstract
We present Generative Relightable Avatars (GRA), a personspecific method for photorealistic free-view rendering and environmentmap relighting of full-body humans. We postulate that modeling finegrained appearance details is inherently a one-to-many problem that can benefit from a generative formulation. In contrast to fully regressive relightable avatar methods, GRA follows a hybrid approach that combines controllable, physics-grounded relighting with probabilistic refinement. Starting from a tracked animated mesh, we optimize material parameters in UV-space and render a coarse relit appearance under a target HDR environment map. Next, we refine the textures with a feedforward model to capture pose-dependent texture dynamics and illumination effects beyond simplified reflectance assumptions. Finally, a finetuned video-to-video diffusion model transforms the physically grounded renderings into temporally coherent, high-detail videos while preserving 3D control, with an error-recycling strategy for generating long videos. Experimental evaluations demonstrate our method’s improved perceptual quality over prior relightable avatar baselines. We urge the readers to watch the supplementary video. See the project page for more details.