G$^{3}$AN: Disentangling Appearance and Motion for Video Generation

Wang, Yaohui; Bilinski, Piotr; Bremond, Francois; Dantcheva, Antitza

Computer Science > Computer Vision and Pattern Recognition

arXiv:1912.05523v2 (cs)

[Submitted on 11 Dec 2019 (v1), revised 30 Mar 2020 (this version, v2), latest version 13 Jun 2020 (v3)]

Title:G$^{3}$AN: Disentangling Appearance and Motion for Video Generation

Authors:Yaohui Wang, Piotr Bilinski, Francois Bremond, Antitza Dantcheva

View PDF

Abstract:Creating realistic human videos entails the challenge of being able to simultaneously generate both appearance, as well as motion. To tackle this challenge, we introduce G$^{3}$AN, a novel spatio-temporal generative model, which seeks to capture the distribution of high dimensional video data and to model appearance and motion in disentangled manner. The latter is achieved by decomposing appearance and motion in a three-stream Generator, where the main stream aims to model spatio-temporal consistency, whereas the two auxiliary streams augment the main stream with multi-scale appearance and motion features, respectively. An extensive quantitative and qualitative analysis shows that our model systematically and significantly outperforms state-of-the-art methods on the facial expression datasets MUG and UvA-NEMO, as well as the Weizmann and UCF101 datasets on human action. Additional analysis on the learned latent representations confirms the successful decomposition of appearance and motion. Source code and pre-trained models are publicly available.

Comments:	CVPR 2020
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1912.05523 [cs.CV]
	(or arXiv:1912.05523v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1912.05523

Submission history

From: Yaohui Wang [view email]
[v1] Wed, 11 Dec 2019 18:46:53 UTC (8,199 KB)
[v2] Mon, 30 Mar 2020 17:56:54 UTC (8,198 KB)
[v3] Sat, 13 Jun 2020 10:58:55 UTC (8,198 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:G$^{3}$AN: Disentangling Appearance and Motion for Video Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:G$^{3}$AN: Disentangling Appearance and Motion for Video Generation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators