Accelerating Neural Networks with Model Compression Techniques
-
DOI:
https://doi.org/10.67228/30713498/IJADSMC-2023PI0Q6RPublished 01-04-2023
Model Compression, Neural Network Acceleration, Pruning, Quantization, Knowledge Distillation, Edge AI, Efficient Deep Learning Issue
Section
ArticlesHow to Cite
[1]D. Rodríguez, “Accelerating Neural Networks with Model Compression Techniques”, IJADSMC, vol. 6, no. 1, pp. 01–14, Jan. 2023, doi: 10.67228/30713498/IJADSMC-2023PI0Q6R.Abstract
Deep neural networks (DNNs) have achieved outstanding performance in areas such as computer vision, speech recognition, natural language processing, and autonomous systems. However, their high computational cost, memory usage, and energy consumption limit deployment in resource-constrained environments like mobile and edge devices. Model compression has emerged as a crucial solution to improve efficiency while maintaining accuracy. This paper provides a comprehensive study of neural network compression techniques, including pruning, quantization, low-rank factorization, knowledge distillation, and neural architecture optimization. These methods are analyzed based on compression ratio, latency, memory efficiency, and accuracy trade-offs. The study also explores hybrid compression approaches and proposes a systematic workflow from model training to deployment on constrained hardware. Experimental results demonstrate that effective compression significantly reduces model size and computational cost with minimal performance loss. The paper highlights the importance of compression-aware design and concludes as a valuable reference for building efficient and scalable AI systems.
References
[1] Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” in Advances in Neural Information Processing Systems, vol. 2, pp. 598–605, 1990.
[2] B. Hassibi and D. G. Stork, “Second order derivatives for network pruning: Optimal brain surgeon,” in Advances in Neural Information Processing Systems, vol. 5, pp. 164–171, 1993.
[3] S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural networks,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), pp. 1135–1143, 2015.
[4] S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding,” in Proc. Int. Conf. Learning Representations (ICLR), 2016.
[5] T. Courbariaux, Y. Bengio, and J.-P. David, “BinaryConnect: Training deep neural networks with binary weights during propagations,” in Advances in Neural Information Processing Systems, vol. 28, pp. 3123–3131, 2015.
[6] M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “XNOR-Net: ImageNet classification using binary convolutional neural networks,” in Proc. European Conf. Computer Vision (ECCV), pp. 525–542, 2016.
[7] Y. Choi, M. El-Khamy, and J. Lee, “Towards the limit of network quantization,” in Proc. Int. Conf. Learning Representations (ICLR), 2017.
[8] A. Zhou, A. Yao, Y. Guo, L. Xu, and Y. Chen, “Incremental network quantization: Towards lossless CNNs with low-precision weights,” in Proc. Int. Conf. Learning Representations (ICLR), 2017.
[9] W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in Advances in Neural Information Processing Systems, vol. 29, pp. 2074–2082, 2016.
[10] J. Frankle and M. Carbin, “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” in Proc. Int. Conf. Learning Representations (ICLR), 2019.
[11] G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015.
[12] A. Romero et al., “FitNets: Hints for thin deep nets,” in Proc. Int. Conf. Learning Representations (ICLR), 2015.
[13] S. Zagoruyko and N. Komodakis, “Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer,” in Proc. Int. Conf. Learning Representations (ICLR), 2017.
[14] H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in Proc. Int. Conf. Learning Representations (ICLR), 2017.
[15] Y. He, X. Zhang, and J. Sun, “Channel pruning for accelerating very deep neural networks,” in Proc. IEEE Int. Conf. Computer Vision (ICCV), pp. 1389–1397, 2017.
Downloads
How to Cite
[1]D. Rodríguez, “Accelerating Neural Networks with Model Compression Techniques”, IJADSMC, vol. 6, no. 1, pp. 01–14, Jan. 2023, doi: 10.67228/30713498/IJADSMC-2023PI0Q6R.