Batch Kalman Normalization: Towards Training Deep Neural Networks with Micro-Batches

Wang, Guangrun; Peng, Jiefeng; Luo, Ping; Wang, Xinjiang; Lin, Liang

Computer Science > Computer Vision and Pattern Recognition

arXiv:1802.03133 (cs)

[Submitted on 9 Feb 2018 (v1), last revised 28 Feb 2018 (this version, v2)]

Title:Batch Kalman Normalization: Towards Training Deep Neural Networks with Micro-Batches

Authors:Guangrun Wang, Jiefeng Peng, Ping Luo, Xinjiang Wang, Liang Lin

View PDF

Abstract:As an indispensable component, Batch Normalization (BN) has successfully improved the training of deep neural networks (DNNs) with mini-batches, by normalizing the distribution of the internal representation for each hidden layer. However, the effectiveness of BN would diminish with scenario of micro-batch (e.g., less than 10 samples in a mini-batch), since the estimated statistics in a mini-batch are not reliable with insufficient samples. In this paper, we present a novel normalization method, called Batch Kalman Normalization (BKN), for improving and accelerating the training of DNNs, particularly under the context of micro-batches. Specifically, unlike the existing solutions treating each hidden layer as an isolated system, BKN treats all the layers in a network as a whole system, and estimates the statistics of a certain layer by considering the distributions of all its preceding layers, mimicking the merits of Kalman Filtering. BKN has two appealing properties. First, it enables more stable training and faster convergence compared to previous works. Second, training DNNs using BKN performs substantially better than those using BN and its variants, especially when very small mini-batches are presented. On the image classification benchmark of ImageNet, using BKN powered networks we improve upon the best-published model-zoo results: reaching 74.0% top-1 val accuracy for InceptionV2. More importantly, using BKN achieves the comparable accuracy with extremely smaller batch size, such as 64 times smaller on CIFAR-10/100 and 8 times smaller on ImageNet.

Comments:	We presented how to improve and accelerate the training of DNNs, particularly under the context of micro-batches. (Submitted to IJCAI 2018)
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1802.03133 [cs.CV]
	(or arXiv:1802.03133v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1802.03133

Submission history

From: Liang Lin [view email]
[v1] Fri, 9 Feb 2018 05:19:16 UTC (384 KB)
[v2] Wed, 28 Feb 2018 02:01:50 UTC (384 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Batch Kalman Normalization: Towards Training Deep Neural Networks with Micro-Batches

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Batch Kalman Normalization: Towards Training Deep Neural Networks with Micro-Batches

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators