Self-Healing Data Pipelines for Fault-Tolerant Data Processing
-
DOI:
https://doi.org/10.67228/30715717/IJDEIC-2025PI4R9APublished 06-05-2025
Self-Healing Data Pipelines, Fault-Tolerant Systems, Distributed Data Processing, Autonomous Recovery, Machine Learning, Intelligent Monitoring, Data Engineering, Workflow Orchestration, Predictive Maintenance, Distributed Computing, Stream Processing, Cloud-Native Architectures Issue
Section
ArticlesHow to Cite
[1]N. Karmarkar, “Self-Healing Data Pipelines for Fault-Tolerant Data Processing”, IJDEIC, vol. 8, no. 1, pp. 01–17, Jun. 2025, doi: 10.67228/30715717/IJDEIC-2025PI4R9A.Abstract
Modern enterprise data ecosystems process information from diverse sources such as cloud applications, IoT devices, distributed databases, and social media. Traditional data pipelines often depend on manual fault handling and static recovery methods, leading to downtime, data inconsistency, and operational inefficiencies. Self-healing data pipelines address these challenges through autonomous fault detection, diagnosis, and recovery using AI, machine learning, anomaly detection, workflow orchestration, and predictive analytics. Techniques such as automated retry, dynamic rerouting, checkpoint recovery, adaptive scheduling, and intelligent resource allocation improve system resilience and fault tolerance. Technologies including Apache Kafka, Apache Spark, Apache Flink, and Kubernetes support scalable and autonomous pipeline management. This study proposes a framework for anomaly detection, root cause analysis, automatic remediation, and adaptive optimization in distributed data systems. Experimental results show improved availability, throughput, recovery efficiency, and reduced Mean Time to Recovery (MTTR), demonstrating that AI-driven self-healing pipelines significantly enhance scalability, resilience, and operational continuity in enterprise analytics infrastructures.
References
[1] J. Gray and A. Reuter, Transaction Processing: Concepts and Techniques. San Francisco, CA, USA: Morgan Kaufmann, 1993.
[2] M. Stonebraker, “The case for shared nothing,” IEEE Database Engineering Bulletin, vol. 9, no. 1, pp. 4–9, 1986.
[3] J. Dean and S. Ghemawat, “MapReduce: Simplified data processing on large clusters,” in Proc. 6th USENIX Symp. Operating Systems Design and Implementation (OSDI), San Francisco, CA, USA, 2004, pp. 137–150.
[4] T. White, Hadoop: The Definitive Guide, 4th ed. Sebastopol, CA, USA: O’Reilly Media, 2015.
[5] M. Zaharia et al., “Spark: Cluster computing with working sets,” in Proc. 2nd USENIX Conf. Hot Topics in Cloud Computing (HotCloud), Boston, MA, USA, 2010, pp. 1–7.
[6] P. Carbone et al., “Apache Flink: Stream and batch processing in a single engine,” IEEE Data Engineering Bulletin, vol. 38, no. 4, pp. 28–38, 2015.
[7] N. Marz and J. Warren, Big Data: Principles and Best Practices of Scalable Realtime Data Systems. Greenwich, CT, USA: Manning Publications, 2015.
[8] G. DeCandia et al., “Dynamo: Amazon’s highly available key-value store,” in Proc. 21st ACM Symp. Operating Systems Principles (SOSP), Stevenson, WA, USA, 2007, pp. 205–220.
[9] J. Kreps, N. Narkhede, and J. Rao, “Kafka: A distributed messaging system for log processing,” in Proc. NetDB Workshop, Athens, Greece, 2011, pp. 1–7.
[10] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.
[11] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
[12] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016, pp. 785–794.
[13] J. O. Kephart and D. M. Chess, “The vision of autonomic computing,” Computer, vol. 36, no. 1, pp. 41–50, Jan. 2003.
[14] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, “Borg, Omega, and Kubernetes,” Communications of the ACM, vol. 59, no. 5, pp. 50–57, 2016.
[15] R. Buyya, C. S. Yeo, and S. Venugopal, “Market-oriented cloud computing: Vision, hype, and reality for delivering IT services as computing utilities,” in Proc. 10th IEEE Int. Conf. High Performance Computing and Communications, Dalian, China, 2008, pp. 5–13.
[16] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.
[17] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845
[18] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820
Downloads
How to Cite
[1]N. Karmarkar, “Self-Healing Data Pipelines for Fault-Tolerant Data Processing”, IJDEIC, vol. 8, no. 1, pp. 01–17, Jun. 2025, doi: 10.67228/30715717/IJDEIC-2025PI4R9A.