Explainable AI Techniques for Intelligent Data Pipeline Optimization
-
DOI:
https://doi.org/10.67228/30715717/IJDEIC-2025PII7D2PPublished 09-04-2025
Explainable Artificial Intelligence (Xai), Intelligent Data Pipelines, Data Pipeline Optimization, Machine Learning, Workflow Scheduling, Predictive Analytics, Feature Attribution, Resource Optimization, Data Engineering, Trustworthy Ai, Pipeline Automation, Decision Support Systems Issue
Section
ArticlesHow to Cite
[1]P. B. Hansen and O. S. Olesen, “Explainable AI Techniques for Intelligent Data Pipeline Optimization”, IJDEIC, vol. 8, no. 2, pp. 01–14, Sep. 2025, doi: 10.67228/30715717/IJDEIC-2025PII7D2P.Abstract
Modern enterprises rely on intelligent data pipelines to collect, process, transform, and analyze data from diverse sources such as cloud platforms, IoT devices, enterprise systems, and social media. Traditional optimization techniques, including rule-based scheduling and heuristic resource allocation, improve efficiency but struggle to adapt to dynamic workloads, changing resource availability, and evolving business requirements. Artificial Intelligence (AI) addresses these limitations through predictive analytics, adaptive scheduling, anomaly detection, and autonomous resource optimization. However, the opaque nature of many AI models reduces transparency, trust, and regulatory compliance. This paper proposes an Explainable AI (XAI)-based Intelligent Data Pipeline Optimization Framework that integrates data preprocessing, predictive analytics, explainability, and adaptive optimization. The framework continuously monitors pipeline performance, generates optimization recommendations, and provides human-interpretable explanations for AI-driven decisions using feature attribution and model interpretation techniques. An automated feedback mechanism enables continuous learning and improvement. Experimental evaluation demonstrates enhanced optimization accuracy, reliability, scalability, interpretability, and administrator trust with minimal impact on performance. The proposed framework provides a transparent and trustworthy approach for next-generation intelligent data engineering systems.
References
[1] T. White, Hadoop: The Definitive Guide, 4th ed. Sebastopol, CA, USA: O'Reilly Media, 2015.
[2] M. Zaharia et al., "Apache Spark: A Unified Engine for Big Data Processing," Commun. ACM, vol. 59, no. 11, pp. 56–65, Nov. 2016.
[3] J. Dean and S. Ghemawat, "MapReduce: Simplified Data Processing on Large Clusters," Commun. ACM, vol. 51, no. 1, pp. 107–113, Jan. 2008.
[4] M. Armbrust et al., "Spark SQL: Relational Data Processing in Spark," in Proc. ACM SIGMOD Int. Conf. Manage. Data, 2015, pp. 1383–1394.
[5] F. Pedregosa et al., "Scikit-learn: Machine Learning in Python," J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.
[6] L. Breiman, "Random Forests," Mach. Learn., vol. 45, no. 1, pp. 5–32, Oct. 2001.
[7] C. Cortes and V. Vapnik, "Support-Vector Networks," Mach. Learn., vol. 20, no. 3, pp. 273–297, Sept. 1995.
[8] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.
[9] S. Hochreiter and J. Schmidhuber, "Long Short-Term Memory," Neural Comput., vol. 9, no. 8, pp. 1735–1780, Nov. 1997.
[10] A. Vaswani et al., "Attention Is All You Need," in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 5998–6008.
[11] M. T. Ribeiro, S. Singh, and C. Guestrin, "Why Should I Trust You? Explaining the Predictions of Any Classifier," in Proc. ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 1135–1144.
[12] S. M. Lundberg and S.-I. Lee, "A Unified Approach to Interpreting Model Predictions," in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 4765–4774.
[13] D. Gunning and D. Aha, "DARPA's Explainable Artificial Intelligence (XAI) Program," AI Mag., vol. 40, no. 2, pp. 44–58, Summer 2019.
[14] R. Guidotti et al., "A Survey of Methods for Explaining Black Box Models," ACM Comput. Surveys, vol. 51, no. 5, pp. 1–42, Jan. 2019.
[15] Z. Chen, X. Liu, Y. Wang, and H. Zhang, "Explainable Artificial Intelligence for Intelligent Data Analytics: A Survey," IEEE Access, vol. 10, pp. 70983–71006, 2022.
[16] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.
[17] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845
Downloads
How to Cite
[1]P. B. Hansen and O. S. Olesen, “Explainable AI Techniques for Intelligent Data Pipeline Optimization”, IJDEIC, vol. 8, no. 2, pp. 01–14, Sep. 2025, doi: 10.67228/30715717/IJDEIC-2025PII7D2P.