DiffProxy: Multi-View Human Mesh Recovery via Diffusion-Generated Dense Proxies
Abstract
Precise human mesh recovery (HMR) from multi-view im-ages remains challenging: end-to-end methods produce entangled errorshard to localize, while x001Ctting-based methods rely on sparse keypointsthat provide limited surface constraints. We observe that the true bot-tleneck lies in the quality of intermediate representations, and that densepixel-to-surface correspondences can be ex001Bectively generated by repur-posing pre-trained dix001Busion models with rich visual priors. We proposeDix001BProxy, a Stable-Dix001Busion-based framework trained on large-scalesynthetic data with pixel-perfect annotations. A multi-conditional proxygenerator predicts dense correspondences from multi-view images, pro-viding uniform surface constraints that enable precise x001Ctting. Hand re-x001Cnement feeds enlarged hand crops alongside full-body images for x001Cne-grained detail, while test-time scaling exploits dix001Busion stochasticity toestimate per-pixel uncertainty. Trained only on synthetic data, Dix001BProxyachieves state-of-the-art results on x001Cve diverse real-world benchmarks.Project page: https://wrk226.github.io/DiffProxy.html