Energy-Efficient ML Inference Pipelines for IoT Data Using Serverless Cloud Functions
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2019PII2W73Published 10-04-2019
Serverless Computing, Energy-Efficient Machine Learning, Iot Data Processing, Ml, Inference Pipelines, Event-Driven Architecture, Model Optimization, Cloud Functions (Faas), Edge–Cloud Collaboration, Sustainable Computing, Low-Latency Inference Issue
Section
ArticlesHow to Cite
[1]S. Diallo and G. Ndlovu, “Energy-Efficient ML Inference Pipelines for IoT Data Using Serverless Cloud Functions”, IJAIDT, vol. 2, no. 2, pp. 01–19, Oct. 2019, doi: 10.67228/30713315/IJAIDT-2019PII2W73.Abstract
The rapid proliferation of Internet of Things (IoT) devices has resulted in massive, continuous data generation, demanding scalable, low-latency, and energy-efficient processing methodologies. Traditional cloud-based machine learning (ML) inference pipelines often incur high energy consumption due to persistent server provisioning and inefficient resource utilization. This paper proposes an energy-efficient ML inference framework using serverless cloud functions that dynamically scale with IoT workloads. The architecture leverages event-driven execution, model optimization techniques (quantization, pruning, edge pre-filtering), and adaptive model selection based on workload intensity. Experimental evaluations conducted on widely used serverless platforms demonstrate significant reductions in energy consumption, cold-start latency, and operational cost while maintaining high inference accuracy. The study highlights the potential of serverless computing as a sustainable backbone for next-generation IoT–ML systems, offering guidelines for building carbon-aware and cost-efficient inference pipelines for real-world applications.
References
[1] S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization, and Huffman coding,” Proc. ICLR, 2016.
[2] G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015.
[3] Satyanarayanan, M. (2017). The emergence of edge computing. Computer, 50(1), 30–39.
[4] Shi, W., Cao, J., Zhang, Q., Li, Y., & Xu, L. (2016). Edge computing: Vision and challenges. IEEE Internet of Things Journal, 3(5), 637–646.
[5] Mao, Y., Zhang, J., & Letaief, K. B. (2017). Dynamic computation offloading for mobile-edge computing with energy harvesting devices. IEEE Journal on Selected Areas in Communications, 34(12), 3590–3605.
[6] Sardellitti, S., Scutari, G., & Barbarossa, S. (2015). Joint optimization of radio and computational resources for multicell mobile-edge computing. IEEE Transactions on Signal and Information Processing over Networks, 1(2), 89–103.
[7] Deng, R., Lu, R., Lai, C., Luan, T. H., & Liang, H. (2016). Optimal workload allocation in fog-cloud computing toward balanced delay and power consumption. IEEE Internet of Things Journal, 3(6), 1171–1181.
[8] Hellerstein, J. M., Faleiro, J., Gonzalez, J., Schleier-Smith, J., Sreekanti, V., Tumanov, A., & Wu, C. (2018). Serverless computing: One step forward, two steps back. CIDR.
[9] Jonas, E., Schleier-Smith, J., Sreekanti, V., Tsai, C.-C., Khandelwal, A., Pu, Q., … Stoica, I. (2017). Cloud programming simplified: A Berkeley view on serverless computing. arXiv preprint arXiv:1902.03383.
[10] McGrath, G., & Brenner, P. R. (2017). Serverless computing: Design, implementation, and performance. 2017 IEEE 37th International Conference on Distributed Computing Systems Workshops (ICDCSW), 405–410.
[11] Ao, L., Izhikevich, L., Voelker, G. M., & Porter, G. (2018). Sprocket: A serverless video processing framework. Proceedings of the ACM Symposium on Cloud Computing, 263–274.
[12] Ishakian, V., Muthusamy, V., & Slominski, A. (2018). Serving deep learning models in a serverless platform. IEEE International Conference on Cloud Engineering (IC2E), 257–262.
[13] Carreira, J., Fonseca, P., Tumanov, A., Zhang, A., & Katz, R. (2018). A case for serverless machine learning. Workshop on Systems for ML (SysML).
[14] Zhang, Q., Chen, M., & Li, L. (2017). Deep learning-based energy-efficient resource management for IoT systems. IEEE Network, 31(5), 70–76.
[15] Xu, X., Chen, Y., Li, Q., Liu, Y., & Zhang, W. (2018). Efficient multi-user computation offloading for mobile-edge cloud computing. IEEE/ACM Transactions on Networking, 26(3), 1485–1498.
Downloads
How to Cite
[1]S. Diallo and G. Ndlovu, “Energy-Efficient ML Inference Pipelines for IoT Data Using Serverless Cloud Functions”, IJAIDT, vol. 2, no. 2, pp. 01–19, Oct. 2019, doi: 10.67228/30713315/IJAIDT-2019PII2W73.