Autonomous Data Pipeline Optimization Using Reinforcement Learning Techniques

  • Authors

    • Mr. Rahul Mehta Senior Software, Engineer Wipro Ltd, India. Author

    DOI:

    https://doi.org/10.67228/30713498/IJADSMC-2025PI5C8Q

    Published 06-04-2025

  • Autonomous Data Pipelines, Reinforcement Learning, Data Engineering, Pipeline Optimization, Deep Q-Network, Proximal Policy Optimization, Intelligent Scheduling, Resource Allocation, Cloud Computing, Self-Adaptive Systems

    Issue

    Section

    Articles

    How to Cite

    [1]
    R. Mehta, “Autonomous Data Pipeline Optimization Using Reinforcement Learning Techniques”, IJADSMC, vol. 8, no. 1, pp. 01–17, Jun. 2025, doi: 10.67228/30713498/IJADSMC-2025PI5C8Q.
  • Abstract

    Modern large-scale data pipelines support analytics, AI, ML, and real-time applications but face challenges related to scalability, resource utilization, reliability, and changing workloads. This paper proposes a reinforcement learning (RL)-based autonomous optimization framework that integrates RL agents with data orchestration platforms to continuously monitor pipeline states and optimize operations. The framework uses system metrics such as workload patterns, queue lengths, execution delays, resource consumption, and failure rates to make intelligent decisions on task scheduling, resource allocation, workload balancing, and fault recovery. Three RL algorithms—Q-learning, Deep Q-Networks (DQN), and Proximal Policy Optimization (PPO)—are evaluated. Experimental results demonstrate improved throughput, reduced latency, enhanced fault tolerance, and better resource efficiency compared to traditional rule-based approaches. The proposed framework enables adaptive, self-managing data pipelines that improve scalability, resilience, and operational efficiency across enterprise, cloud, and edge environments.

  • References

    [1] R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed. Cambridge, MA, USA: MIT Press, 2018.

    [2] V. Mnih, K. Kavukcuoglu, D. Silver, et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015.

    [3] J. Dean and S. Ghemawat, “MapReduce: Simplified data processing on large clusters,” Communications of the ACM, vol. 51, no. 1, pp. 107–113, 2008.

    [4] M. Zaharia, M. Chowdhury, T. Das, et al., “Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing,” in Proc. USENIX NSDI, 2012, pp. 15–28.

    [5] T. Akidau, A. Balikov, K. Bekiroğlu, et al., “The Dataflow Model: A practical approach to balancing correctness, latency, and cost in massive-scale systems,” Proceedings of the VLDB Endowment, vol. 8, no. 12, pp. 1792–1803, 2015.

    [6] P. Carbone, G. Katsifodimos, S. Ewen, et al., “Apache Flink: Stream and batch processing in a single engine,” IEEE Data Engineering Bulletin, vol. 38, no. 4, pp. 28–38, 2015.

    [7] C. Delimitrou and C. Kozyrakis, “Quasar: Resource-efficient and QoS-aware cluster management,” in Proc. ASPLOS, 2014, pp. 127–144.

    [8] H. Mao, M. Alizadeh, I. Menache, and S. Kandula, “Resource management with deep reinforcement learning,” in Proc. ACM HotNets, 2016, pp. 50–56.

    [9] H. Mao, M. Schwarzkopf, S. Venkatakrishnan, et al., “Learning scheduling algorithms for data processing clusters,” in Proc. ACM SIGCOMM, 2019, pp. 270–288.

    [10] T. Chen, Y. Zhang, and X. Wang, “Deep reinforcement learning for resource allocation in cloud computing environments,” Future Generation Computer Systems, vol. 95, pp. 395–408, 2019.

    [11] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” arXiv:1707.06347, 2017.

    [12] T. P. Lillicrap, J. J. Hunt, A. Pritzel, et al., “Continuous control with deep reinforcement learning,” arXiv:1509.02971, 2015.

    [13] V. R. Konda and J. N. Tsitsiklis, “Actor-Critic algorithms,” in Advances in Neural Information Processing Systems (NIPS), vol. 12, pp. 1008–1014, 2000.

    [14] A. Verma, L. Cherkasova, and R. H. Campbell, “ARIA: Automatic resource inference and allocation for MapReduce environments,” in Proc. ACM ICAC, 2011, pp. 235–244.

    [15] Y. Liu, M. Peng, Y. Shou, Y. Chen, and S. Chen, “Deep reinforcement learning based intelligent resource management for cloud computing,” IEEE Access, vol. 8, pp. 77537–77547, 2020.

    [16] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820.

    [17] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845.

    [18] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.

  • Downloads