JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video
Abstract
Generating realistic human avatars in complex motions—suchas clothing dynamics—requires modeling of global and local deformationswhich remains challenging in monocular settings. We address this prob-lem by leveraging neural Jacobian fields (NJFs) for representing semi-rigid deformations. We train self-supervised neural networks for predict-ing Jacobian matrices that give the pose-dependent deformations, bysolving a Poisson equation. However, monocular input presents severaldifficulties such as self-occluded regions and invisible surfaces. To addressthese issues, we introduce three key components: a constrained Poissonsolver, signed distance-based Jacobian regularization, and a deformation-guided residual flow loss, which together suppress boundary artifacts, re-cover frequently occluded regions such as armpits and thighs, and enforcetemporal consistency during motion. Experiments on benchmark and in-the-wild videos demonstrate that our method generates temporally sta-ble and geometrically coherent avatars, outperforming state-of-the-artapproaches.