Optimizing Mesh Animation from Video via Shape Flow Guidance
Abstract
Optimizing vertex deformations from video for mesh anima-tion is constrained by rendering-based reconstruction losses. While ex-isting approaches improve mesh representations, supervision signals, oranimation paradigms, their supervision remains confined to the 2D do-main. Such 2D supervision is limited: it provides no signal for occludedregions and only indirect cues for visible areas. Consequently, these meth-ods often suffer from severe shape and motion artifacts. To address thislimitation, we propose Shape Flow Guidance (SFG), a sequence of 3Dshapes derived from videos, which serves as explicit 3D supervision formesh animation. Specifically, SFG is elicited by intervening in the sam-pling process of a pretrained mesh generator in a training-free manner.We further tailor a skeletal animation model that separates local defor-mation from global transformation. This model enables SFG to guidecomplex local motion while reserving rendering-based losses for simpleglobal motion. Extensive experiments confirm that our method signifi-cantly outperforms prior work qualitatively, quantitatively, and in termsof processing speed. Qualitative results are available on our project page:https://sfgmesh.pages.dev/.