Integration of DataOps with Automated Exploratory Data Analysis (EDA) Workflows

  • Authors

    • Dana Scott Professor, Emeritus Carnegie Mellon University, United States. Author
    • John Cocke IBM Fellow, IBM Research, United States. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2023PI5R8W

    Published 04-12-2023

  • DataOps, Automated EDA, Exploratory Data Analysis, Data Pipelines, CI/CD in Data Science, Data Engineering, AutoEDA Tools, Data Automation, Machine Learning Operations, Data Quality

    Issue

    Section

    Articles

    How to Cite

    [1]
    D. Scott and J. Cocke, “Integration of DataOps with Automated Exploratory Data Analysis (EDA) Workflows”, IJDEIC, vol. 6, no. 1, pp. 01–10, Apr. 2023, doi: 10.67228/30715717/IJDEIC-2023PI5R8W.
  • Abstract

    In the age of data-driven innovation, rapid and reliable insight generation is critical. Exploratory Data Analysis (EDA) plays a foundational role in understanding data quality, structure, and relationships. However, traditional EDA processes are time-consuming and often disconnected from the operationalized data pipelines managed by DataOps. This paper explores the integration of DataOps principles with automated EDA workflows, aiming to enhance collaboration, reproducibility, and scalability in modern data ecosystems. By aligning automation tools with robust CI/CD pipelines, teams can streamline data profiling, visualization, and anomaly detection in real time. We propose an architectural framework for this integration, illustrate a practical use case, and discuss the implications for both data engineering and data science teams. This convergence promises a more agile and intelligent approach to data exploration, ensuring faster time-to-insight and improved data trustworthiness.

  • References

    [1] Breck, E., Polyzotis, N., Roy, S., et al. (2017). The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. Google Research.

    [2] Ertel, W., & Smith, D. (2020). DataOps: The future of data engineering. Journal of Data Science and Analytics, 5(3), 122–137.

    [3] Wickham, H. (2014). Tidy Data. Journal of Statistical Software, 59(10), 1–23.

    [4] Sculley, D., Holt, G., Golovin, D., et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS.

    [5] Microsoft Azure. (2021). MLOps Maturity Model. Retrieved from: https://azure.microsoft.com/en-us/resources/mlops/

    [6] Munzner, T. (2014). Visualization Analysis and Design. CRC Press.

    [7] Ashcraft, L., & Malakar, A. (2021). Kedro: A Framework for Reproducible, Maintainable Data Science Code. QuantumBlack.

    [8] Ghoting, A., Krishnamurthy, R., Pednault, E., et al. (2015). SystemML: Declarative Machine Learning on Spark. VLDB.

    [9] Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2010). Spark: Cluster Computing with Working Sets. HotCloud.

    [10] Alluvium. (2020). Operationalizing Data Science with DataOps. Alluvium Whitepaper.

  • Downloads