Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

Cartillier, Vincent; Ren, Zhile; Jain, Neha; Lee, Stefan; Essa, Irfan; Batra, Dhruv

Computer Science > Computer Vision and Pattern Recognition

arXiv:2010.01191 (cs)

[Submitted on 2 Oct 2020 (v1), last revised 11 Mar 2021 (this version, v3)]

Title:Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

Authors:Vincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee, Irfan Essa, Dhruv Batra

View PDF

Abstract:We study the task of semantic mapping - specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map ("what is where?") from egocentric observations of an RGB-D camera with known pose (via localization sensors). Towards this goal, we present SemanticMapNet (SMNet), which consists of: (1) an Egocentric Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length x width x feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01-16.81% (absolute) on mean-IoU and 3.81-19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the neural episodic memories and spatio-semantic allocentric representations build by SMNet for subsequent tasks in the same space - navigating to objects seen during the tour("Find chair") or answering questions about the space ("How many chairs did you see in the house?"). Project page: this https URL.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2010.01191 [cs.CV]
	(or arXiv:2010.01191v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2010.01191

Submission history

From: Vincent Cartillier [view email]
[v1] Fri, 2 Oct 2020 20:44:46 UTC (12,332 KB)
[v2] Thu, 4 Mar 2021 05:10:53 UTC (12,330 KB)
[v3] Thu, 11 Mar 2021 00:26:51 UTC (12,332 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators