On-Device Training Under 256KB Memory

Lin, Ji; Zhu, Ligeng; Chen, Wei-Ming; Wang, Wei-Chen; Gan, Chuang; Han, Song

Computer Science > Computer Vision and Pattern Recognition

arXiv:2206.15472 (cs)

[Submitted on 30 Jun 2022 (v1), last revised 3 Apr 2024 (this version, v4)]

Title:On-Device Training Under 256KB Memory

Authors:Ji Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang, Chuang Gan, Song Han

View PDF HTML (experimental)

Abstract:On-device training enables the model to adapt to new data collected from the sensors by fine-tuning a pre-trained model. Users can benefit from customized AI models without having to transfer the data to the cloud, protecting the privacy. However, the training memory consumption is prohibitive for IoT devices that have tiny memory resources. We propose an algorithm-system co-design framework to make on-device training possible with only 256KB of memory. On-device training faces two unique challenges: (1) the quantized graphs of neural networks are hard to optimize due to low bit-precision and the lack of normalization; (2) the limited hardware resource does not allow full back-propagation. To cope with the optimization difficulty, we propose Quantization-Aware Scaling to calibrate the gradient scales and stabilize 8-bit quantized training. To reduce the memory footprint, we propose Sparse Update to skip the gradient computation of less important layers and sub-tensors. The algorithm innovation is implemented by a lightweight training system, Tiny Training Engine, which prunes the backward computation graph to support sparse updates and offload the runtime auto-differentiation to compile time. Our framework is the first solution to enable tiny on-device training of convolutional neural networks under 256KB SRAM and 1MB Flash without auxiliary memory, using less than 1/1000 of the memory of PyTorch and TensorFlow while matching the accuracy on tinyML application VWW. Our study enables IoT devices not only to perform inference but also to continuously adapt to new data for on-device lifelong learning. A video demo can be found here: this https URL.

Comments:	NeurIPS 2022
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2206.15472 [cs.CV]
	(or arXiv:2206.15472v4 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2206.15472

Submission history

From: Ji Lin [view email]
[v1] Thu, 30 Jun 2022 17:59:08 UTC (2,108 KB)
[v2] Thu, 14 Jul 2022 03:20:34 UTC (2,108 KB)
[v3] Thu, 10 Nov 2022 18:10:18 UTC (4,929 KB)
[v4] Wed, 3 Apr 2024 03:15:55 UTC (3,054 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:On-Device Training Under 256KB Memory

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:On-Device Training Under 256KB Memory

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators