Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation

Joo, Hanbyul; Neverova, Natalia; Vedaldi, Andrea

Computer Science > Computer Vision and Pattern Recognition

arXiv:2004.03686 (cs)

[Submitted on 7 Apr 2020 (v1), last revised 22 Oct 2021 (this version, v3)]

Title:Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation

Authors:Hanbyul Joo, Natalia Neverova, Andrea Vedaldi

View PDF

Abstract:Differently from 2D image datasets such as COCO, large-scale human datasets with 3D ground-truth annotations are very difficult to obtain in the wild. In this paper, we address this problem by augmenting existing 2D datasets with high-quality 3D pose fits. Remarkably, the resulting annotations are sufficient to train from scratch 3D pose regressor networks that outperform the current state-of-the-art on in-the-wild benchmarks such as 3DPW. Additionally, training on our augmented data is straightforward as it does not require to mix multiple and incompatible 2D and 3D datasets or to use complicated network architectures and training procedures. This simplified pipeline affords additional improvements, including injecting extreme crop augmentations to better reconstruct highly truncated people, and incorporating auxiliary inputs to improve 3D pose estimation accuracy. It also reduces the dependency on 3D datasets such as H36M that have restrictive licenses. We also use our method to introduce new benchmarks for the study of real-world challenges such as occlusions, truncations, and rare body poses. In order to obtain such high quality 3D pseudo-annotations, inspired by progress in internal learning, we introduce Exemplar Fine-Tuning (EFT). EFT combines the re-projection accuracy of fitting methods like SMPLify with a 3D pose prior implicitly captured by a pre-trained 3D pose regressor network. We show that EFT produces 3D annotations that result in better downstream performance and are qualitatively preferable in an extensive human-based assessment.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2004.03686 [cs.CV]
	(or arXiv:2004.03686v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2004.03686

Submission history

From: Hanbyul Joo [view email]
[v1] Tue, 7 Apr 2020 20:21:18 UTC (6,781 KB)
[v2] Wed, 2 Sep 2020 23:05:36 UTC (34,596 KB)
[v3] Fri, 22 Oct 2021 02:55:04 UTC (22,668 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators