FeatER: An Efficient Network for Human Reconstruction via Feature Map-Based TransformER

Zheng, Ce; Mendieta, Matias; Yang, Taojiannan; Qi, Guo-Jun; Chen, Chen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2205.15448 (cs)

[Submitted on 30 May 2022 (v1), last revised 23 Mar 2023 (this version, v3)]

Title:FeatER: An Efficient Network for Human Reconstruction via Feature Map-Based TransformER

Authors:Ce Zheng, Matias Mendieta, Taojiannan Yang, Guo-Jun Qi, Chen Chen

View PDF

Abstract:Recently, vision transformers have shown great success in a set of human reconstruction tasks such as 2D human pose estimation (2D HPE), 3D human pose estimation (3D HPE), and human mesh reconstruction (HMR) tasks. In these tasks, feature map representations of the human structural information are often extracted first from the image by a CNN (such as HRNet), and then further processed by transformer to predict the heatmaps (encodes each joint's location into a feature map with a Gaussian distribution) for HPE or HMR. However, existing transformer architectures are not able to process these feature map inputs directly, forcing an unnatural flattening of the location-sensitive human structural information. Furthermore, much of the performance benefit in recent HPE and HMR methods has come at the cost of ever-increasing computation and memory needs. Therefore, to simultaneously address these problems, we propose FeatER, a novel transformer design that preserves the inherent structure of feature map representations when modeling attention while reducing memory and computational costs. Taking advantage of FeatER, we build an efficient network for a set of human reconstruction tasks including 2D HPE, 3D HPE, and HMR. A feature map reconstruction module is applied to improve the performance of the estimated human pose and mesh. Extensive experiments demonstrate the effectiveness of FeatER on various human pose and mesh datasets. For instance, FeatER outperforms the SOTA method MeshGraphormer by requiring 5% of Params and 16% of MACs on Human3.6M and 3DPW datasets. The project webpage is this https URL.

Comments:	CVPR 2023
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
Cite as:	arXiv:2205.15448 [cs.CV]
	(or arXiv:2205.15448v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2205.15448

Submission history

From: Ce Zheng [view email]
[v1] Mon, 30 May 2022 22:09:57 UTC (7,779 KB)
[v2] Wed, 23 Nov 2022 00:03:20 UTC (13,270 KB)
[v3] Thu, 23 Mar 2023 15:48:05 UTC (13,456 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:FeatER: An Efficient Network for Human Reconstruction via Feature Map-Based TransformER

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:FeatER: An Efficient Network for Human Reconstruction via Feature Map-Based TransformER

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators