AI-Driven Software Engineering: Optimizing Distributed Systems for Scalable Machine Learning Workflows
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2024PI1J5B8Published 06-22-2026
AI-Driven Software Engineering, Distributed Systems, Machine Learning Workflows, Resource Optimization, Intelligent Scheduling, Auto-scaling, Fault Tolerance, Performance Optimization, Cloud Computing, Data Parallelism Issue
Section
ArticlesHow to Cite
[1]Y. Tanaka, “AI-Driven Software Engineering: Optimizing Distributed Systems for Scalable Machine Learning Workflows”, IJAIDT, vol. 7, no. 1, pp. 01–12, Jun. 2026, doi: 10.67228/30713315/IJAIDT-2024PI1J5B8.Abstract
This paper explores the integration of artificial intelligence techniques into software engineering practices to optimize distributed systems for scalable machine learning (ML) workflows. As ML models grow in complexity and data volume, traditional system design approaches struggle to meet the demands of performance, scalability, and resource efficiency. We propose an AI-driven framework that leverages predictive analytics, automated resource management, and intelligent scheduling to enhance distributed computing environments. The study examines key challenges in distributed ML systems, including data partitioning, workload balancing, fault tolerance, and latency optimization. Through a combination of simulation and real-world case studies, we demonstrate how AI-based optimization strategies improve system throughput, reduce training time, and enhance resource utilization. The results highlight the potential of combining software engineering principles with AI-driven decision-making to build resilient and efficient ML infrastructures. This work contributes a structured approach for designing next-generation distributed systems capable of supporting large-scale machine learning applications.
References
[1] Dean, J., and Ghemawat, S. “MapReduce: Simplified Data Processing on Large Clusters.” Communications of the ACM, vol. 51, no. 1, 2008, pp. 107–113.
[2] Abadi, M., et al. “TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems.” Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2016.
[3] Zaharia, M., et al. “Apache Spark: A Unified Engine for Big Data Processing.” Communications of the ACM, vol. 59, no. 11, 2016, pp. 56–65.
[4] Li, M., et al. “Parameter Server for Distributed Machine Learning.” Journal of Machine Learning Research, vol. 18, no. 285, 2017, pp. 1–41.
[5] Hsueh, C.-B., and Chiang, C.-A. “AI-Driven Optimization in Software Engineering Processes: A Survey.” IEEE Transactions on Software Engineering, vol. 47, no. 11, 2021, pp. 2195–2213.
[6] Goyal, P., et al. “Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
[7] Al-Rfou, R., et al. “Distributed Representations for Scalable Deep Learning.” Proceedings of the 33rd International Conference on Machine Learning (ICML), 2016.
[8] Chen, J., Pan, X., and Zhang, T. “Communication-Efficient Distributed Machine Learning: Algorithms and Systems.” Foundations and Trends in Machine Learning, vol. 13, no. 5–6, 2020.
[9] Amershi, S., et al. “Software Engineering for Machine Learning: A Case Study.” Proceedings of the 41st International Conference on Software Engineering (ICSE), 2019.
[10] Yu, L., and Ma, X. “AI-Augmented DevOps: Intelligent Automation for ML Pipelines.” Journal of Systems and Software, vol. 182, 2021.
[11] Kraska, T. “Building Machine Learning Systems: Architectures, Infrastructure, and Tools.” Proceedings of the VLDB Endowment, vol. 12, no. 12, 2019, pp. 2218–2232.
[12] Banerjee, S., et al. “Adaptive Resource Allocation in Distributed ML Systems Using Reinforcement Learning.” IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 4, 2022.
[13] Xu, Y., and Huang, J. “Optimizing System Performance for AI Workloads Using Predictive Software Engineering.” ACM Transactions on Software Engineering and Methodology, vol. 30, no. 2, 2021.
[14] Verma, A., Pedrosa, L., and Korupolu, M. “Large-Scale Cluster Scheduling for Machine Learning Workloads.” Proceedings of the USENIX Annual Technical Conference (ATC), 2015.
[15] Jordan, M. I., and Mitchell, T. M. “Machine Learning: Trends, Perspectives, and Prospects.” Science, vol. 349, no. 6245, 2015, pp. 255–260.
[16] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820
Downloads
How to Cite
[1]Y. Tanaka, “AI-Driven Software Engineering: Optimizing Distributed Systems for Scalable Machine Learning Workflows”, IJAIDT, vol. 7, no. 1, pp. 01–12, Jun. 2026, doi: 10.67228/30713315/IJAIDT-2024PI1J5B8.