Retraining-Based Iterative Weight Quantization for Deep Neural Networks

Lee, Dongsoo; Kim, Byeongwook

Computer Science > Machine Learning

arXiv:1805.11233 (cs)

[Submitted on 29 May 2018]

Title:Retraining-Based Iterative Weight Quantization for Deep Neural Networks

Authors:Dongsoo Lee, Byeongwook Kim

View PDF

Abstract:Model compression has gained a lot of attention due to its ability to reduce hardware resource requirements significantly while maintaining accuracy of DNNs. Model compression is especially useful for memory-intensive recurrent neural networks because smaller memory footprint is crucial not only for reducing storage requirement but also for fast inference operations. Quantization is known to be an effective model compression method and researchers are interested in minimizing the number of bits to represent parameters. In this work, we introduce an iterative technique to apply quantization, presenting high compression ratio without any modifications to the training algorithm. In the proposed technique, weight quantization is followed by retraining the model with full precision weights. We show that iterative retraining generates new sets of weights which can be quantized with decreasing quantization loss at each iteration. We also show that quantization is efficiently able to leverage pruning, another effective model compression method. Implementation issues on combining the two methods are also addressed. Our experimental results demonstrate that an LSTM model using 1-bit quantized weights is sufficient for PTB dataset without any accuracy degradation while previous methods demand at least 2-4 bits for quantized weights.

Comments:	12 pages, 13 figures, NIPS 2018 (32nd Annual Conference on Neural Information Processing Systems) submission
Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1805.11233 [cs.LG]
	(or arXiv:1805.11233v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1805.11233

Submission history

From: Dongsoo Lee [view email]
[v1] Tue, 29 May 2018 03:36:40 UTC (124 KB)

Computer Science > Machine Learning

Title:Retraining-Based Iterative Weight Quantization for Deep Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Retraining-Based Iterative Weight Quantization for Deep Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators