Topology Reduction in Deep Convolutional Feature Extraction Networks

Wiatowski, Thomas; Grohs, Philipp; Bölcskei, Helmut

Statistics > Machine Learning

arXiv:1707.02711 (stat)

[Submitted on 10 Jul 2017 (v1), last revised 14 Mar 2018 (this version, v2)]

Title:Topology Reduction in Deep Convolutional Feature Extraction Networks

Authors:Thomas Wiatowski, Philipp Grohs, Helmut Bölcskei

View PDF

Abstract:Deep convolutional neural networks (CNNs) used in practice employ potentially hundreds of layers and $10$,$000$s of nodes. Such network sizes entail significant computational complexity due to the large number of convolutions that need to be carried out; in addition, a large number of parameters needs to be learned and stored. Very deep and wide CNNs may therefore not be well suited to applications operating under severe resource constraints as is the case, e.g., in low-power embedded and mobile platforms. This paper aims at understanding the impact of CNN topology, specifically depth and width, on the network's feature extraction capabilities. We address this question for the class of scattering networks that employ either Weyl-Heisenberg filters or wavelets, the modulus non-linearity, and no pooling. The exponential feature map energy decay results in Wiatowski et al., 2017, are generalized to $\mathcal{O}(a^{-N})$, where an arbitrary decay factor $a>1$ can be realized through suitable choice of the Weyl-Heisenberg prototype function or the mother wavelet. We then show how networks of fixed (possibly small) depth $N$ can be designed to guarantee that $((1-\varepsilon)\cdot 100)\%$ of the input signal's energy are contained in the feature vector. Based on the notion of operationally significant nodes, we characterize, partly rigorously and partly heuristically, the topology-reducing effects of (effectively) band-limited input signals, band-limited filters, and feature map symmetries. Finally, for networks based on Weyl-Heisenberg filters, we determine the prototype function bandwidth that minimizes---for fixed network depth $N$---the average number of operationally significant nodes per layer.

Comments:	Corrected errors in arguments on spectral decay of Sobolev functions. Replaced part of the decay results (Sections 5-7) by corresponding statements for effectively band-limited functions
Subjects:	Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Information Theory (cs.IT); Machine Learning (cs.LG); Functional Analysis (math.FA)
Cite as:	arXiv:1707.02711 [stat.ML]
	(or arXiv:1707.02711v2 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.1707.02711
Journal reference:	Proc. of SPIE (Wavelets and Sparsity XVII), San Diego, USA, Vol. 10394, pp. 1039418:1-1039418:12, Aug. 2017, (invited paper)

Submission history

From: Thomas Wiatowski [view email]
[v1] Mon, 10 Jul 2017 06:35:48 UTC (55 KB)
[v2] Wed, 14 Mar 2018 08:59:37 UTC (55 KB)

Statistics > Machine Learning

Title:Topology Reduction in Deep Convolutional Feature Extraction Networks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Topology Reduction in Deep Convolutional Feature Extraction Networks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators