Make It Move: Controllable Image-to-Video Generation with Text Descriptions

Hu, Yaosi; Luo, Chong; Chen, Zhenzhong

Computer Science > Computer Vision and Pattern Recognition

arXiv:2112.02815v1 (cs)

[Submitted on 6 Dec 2021 (this version), latest version 31 Mar 2022 (v2)]

Title:Make It Move: Controllable Image-to-Video Generation with Text Descriptions

Authors:Yaosi Hu, Chong Luo, Zhenzhong Chen

View PDF

Abstract:Generating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video generation (TI2V), is proposed. With both controllable appearance and motion, TI2V aims at generating videos from a static image and a text description. The key challenges of TI2V task lie both in aligning appearance and motion from different modalities, and in handling uncertainty in text descriptions. To address these challenges, we propose a Motion Anchor-based video GEnerator (MAGE) with an innovative motion anchor (MA) structure to store appearance-motion aligned representation. To model the uncertainty and increase the diversity, it further allows the injection of explicit condition and implicit randomness. Through three-dimensional axial transformers, MA is interacted with given image to generate next frames recursively with satisfying controllability and diversity. Accompanying the new task, we build two new video-text paired datasets based on MNIST and CATER for evaluation. Experiments conducted on these datasets verify the effectiveness of MAGE and show appealing potentials of TI2V task. Source code for model and datasets will be available soon.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as:	arXiv:2112.02815 [cs.CV]
	(or arXiv:2112.02815v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2112.02815

Submission history

From: Chong Luo [view email]
[v1] Mon, 6 Dec 2021 07:00:36 UTC (26,106 KB)
[v2] Thu, 31 Mar 2022 05:28:53 UTC (26,108 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Make It Move: Controllable Image-to-Video Generation with Text Descriptions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Make It Move: Controllable Image-to-Video Generation with Text Descriptions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators