Impact of Data Versioning on Longitudinal Analytical Model Performance
-
DOI:
https://doi.org/10.67228/30715717/IJDEIC-2022PII8M5PPublished 12-07-2022
Data Versioning, Longitudinal Analysis, Model Performance Drift, Temporal Data, Reproducibility, Predictive Modeling, Data Management, Time-Series, Machine Learning, Data Lineage Issue
Section
ArticlesHow to Cite
[1]H. Zemanek, “Impact of Data Versioning on Longitudinal Analytical Model Performance”, IJDEIC, vol. 5, no. 2, pp. 01–09, Dec. 2022, doi: 10.67228/30715717/IJDEIC-2022PII8M5P.Abstract
In data-driven environments where datasets evolve over time, the challenge of maintaining consistent and high-performing analytical models becomes increasingly critical. This paper investigates the impact of data versioning on the performance of longitudinal analytical models. We explore how changes in data over time, captured through systematic versioning, influence model accuracy, stability, and generalizability. Using real-world longitudinal datasets and a comparative modeling framework, we assess various data versioning strategies and their effects on predictive performance. Our findings reveal that integrating data versioning not only enhances reproducibility but also enables more robust handling of performance drift over time. This research offers practical insights for data scientists and engineers aiming to maintain the fidelity of analytical systems in dynamic environments.
References
[1] Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In Proceedings of SysML Conference.
[2] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, (pp. 785–794).
[3] Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., ... & Dennison, D. (2015). Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems, (pp. 2503–2511).
[4] Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., ... & Chintala, S. (2019). PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems, 32.
[5] Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., & Zhang, G. (2018). Learning under concept drift: A review. IEEE Transactions on Knowledge and Data Engineering, 31(12), 2346–2363.
[6] Zaharia, M., Das, T., Li, H., Hunter, T., Shenker, S., & Stoica, I. (2013). Discretized streams: Fault-tolerant streaming computation at scale. In Proceedings of the 24th ACM Symposium on Operating Systems Principles (pp. 423–438).
[7] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys (CSUR), 46(4), 1–37.
[8] Kuznetsov, S., & Nivargi, R. (2020). Delta Lake: High-performance ACID table storage over cloud object stores. Data Engineering, 43(1), 45–56.
[9] Sato, R., Lin, Y., Shibata, Y., & Ohsuga, A. (2021). Reproducible machine learning with Pachyderm: Data versioning and pipeline orchestration. In IEEE International Conference on Big Data (pp. 3029–3038).
Downloads
How to Cite
[1]H. Zemanek, “Impact of Data Versioning on Longitudinal Analytical Model Performance”, IJDEIC, vol. 5, no. 2, pp. 01–09, Dec. 2022, doi: 10.67228/30715717/IJDEIC-2022PII8M5P.