Impact of Data Transformation on Model Drift in MLOps Pipelines

  • Authors

    • Dr. R. Kartikeyan Assosciate Professor, Anna University, Chennai, India. Author

    DOI:

    https://doi.org/10.67228/30713315/IJAIDT-2020PII0V9H

    Published 12-05-2020

  • Model Drift, Data Transformation, Mlops, Data Drift, Feature Engineering, Transformation Lineage, Continuous Validation, Machine Learning Pipelines, Concept Drift, Model Monitoring

    Issue

    Section

    Articles

    How to Cite

    [1]
    K. R, “Impact of Data Transformation on Model Drift in MLOps Pipelines”, IJAIDT, vol. 3, no. 2, pp. 01–12, Dec. 2020, doi: 10.67228/30713315/IJAIDT-2020PII0V9H.
  • Abstract

    Data transformation is a critical pre-processing step in machine learning pipelines, shaping how raw data is cleaned, normalized, enriched, and structured for training and inference. However, as data evolves over time, even small, unmonitored changes in transformation logic can lead to significant deviations in model behavior a phenomenon known as model drift. This paper explores the interplay between data transformation and model drift within MLOps pipelines, analyzing how transformation pipelines contribute to both data drift and concept drift. We examine how subtle modifications in schema, feature engineering, or data semantics impact model performance and reliability. Through real-world case studies and reference architectures, we highlight best practices for tracking transformation lineage, validating feature consistency, and implementing drift detection mechanisms that include transformation awareness. Additionally, the paper proposes a framework for making data transformation pipelines "drift-aware" by integrating version control, observability, and automated validation. As MLOps matures, the ability to trace, audit, and govern transformations becomes essential for maintaining model integrity and compliance in production environments.

  • References

    [1] Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., ... & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503-2511.

    [2] Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., ... & Zimmermann, T. (2019). Software engineering for machine learning: A case study. ICSE, 291-300.

    [3] Breck, E., Polyzotis, N., & Whang, S. E. (2019). Data validation for machine learning. CIDR.

    [4] Sato, I., & Nakagawa, S. (2020). Handling data drift in machine learning pipelines. Journal of Data and Information Quality, 12(4), 15.

    [5] Hota, A. R., & Singh, A. (2020). ML governance and pipeline versioning for mitigating data drift. Proceedings of the ACM Symposium on Cloud Computing.

    [6] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. KDD, 785-794.

    [7] Zinkevich, M., & Liang, J. (2020). Continuous training: A framework for ML model lifecycle management.

    [8] Luo, H., Zhou, W., & Liu, Q. (2019). Detecting concept drift in data streams using feature distribution monitoring. Data Mining and Knowledge Discovery, 33(4), 1037-1067.

    [9] Amershi, S., Lee, B., Kapoor, A., & Tan, D. S. (2018). ModelTracker: Redesigning performance analysis tools for machine learning. CHI.

    [10] Rajaraman, A., & Ullman, J. D. (2011). Mining of massive datasets. Cambridge University Press.

    [11] Wang, T., & Abraham, A. (2019). Feature store: Managing data for machine learning. IEEE Data Engineering Bulletin.

    [12] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 1-37.

    [13] Zaharia, M., et al. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56-65.

    [14] Cheng, J., et al. (2020). Model monitoring and drift detection with Evidently AI. Proceedings of the ML Systems Workshop.

  • Downloads