Data-Centric Refactoring: Techniques for Improving Model Quality via Codebase Changes
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2019PI6J4FPublished 04-04-2019
Data-Centric AI, Software Refactoring, Ml Pipelines, Feature Engineering, Technical Debt, Data Quality, Codebase Optimization, Mlops, Model Reliability, Data Preprocessing Issue
Section
ArticlesHow to Cite
[1]F. Diop, “Data-Centric Refactoring: Techniques for Improving Model Quality via Codebase Changes”, IJAIDT, vol. 2, no. 1, pp. 01–15, Apr. 2019, doi: 10.67228/30713315/IJAIDT-2019PI6J4F.Abstract
Machine learning systems often fail to reach optimal performance not because of inadequate model architectures, but due to poorly structured data processing pipelines hidden within the codebase. Data-centric refactoring aims to improve model quality through systematic restructuring of code elements responsible for data collection, preprocessing, transformation, validation, and feature engineering. This paper introduces a comprehensive taxonomy of data-centric refactoring strategies, investigates their application across ML-driven software projects, and evaluates their impact on model accuracy, robustness, maintainability, and reproducibility. By bridging software refactoring principles with data-centric AI practices, the proposed framework demonstrates that code-level improvements to data handling routines can yield substantial gains in model performance while reducing technical debt. Experimental results show that systematically refactoring data pipelines leads to more reliable features, reduced noise propagation, and improved generalization. The findings position data-centric refactoring as a key discipline for modern ML engineering, enabling scalable, interpretable, and production-ready models.
References
[1] Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., & Zimmermann, T. (2019). Software engineering for machine learning: A case study. IEEE Software, 36(5), 50–57.
[2] Breck, E., Cai, S., Nielsen, E., Polyzotis, N., Roy, S., & Whang, S. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. Proceedings of SysML Workshop.
[3] Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., et al. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems (NeurIPS).
[4] Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2018). Data management challenges in production machine learning. ACM SIGMOD Record, 47(1), 17–28.
[5] Kim, M., Zimmermann, T., & Nagappan, N. (2016). A field study of refactoring challenges and benefits. ACM Transactions on Software Engineering and Methodology, 25(2), 1–30.
[6] Palomba, F., & Bavota, G. (2019). Learning-to-refactor: Automated refactoring decision support via ML. ICSE Proceedings, IEEE, 993–1004.
[7] Fowler, M. (2002). Refactoring: Improving the design of existing code. Addison-Wesley.
[8] Opdyke, W. F. (1992). Refactoring object-oriented frameworks (Doctoral dissertation, University of Illinois).
[9] Mens, T., & Tourwé, T. (2004). A survey of software refactoring. IEEE Transactions on Software Engineering, 30(2), 126–139.
[10] Murphy-Hill, E., Parnin, C., & Black, A. P. (2012). How we refactor, and how we know it. IEEE Transactions on Software Engineering, 38(1), 5–18.
[11] Kim, M., Zimmermann, T., & Nagappan, N. (2014). A field study of refactoring challenges and benefits. Proceedings of the ACM SIGSOFT FSE, 50–60.
[12] Bavota, G., De Lucia, A., Di Penta, M., Oliveto, R., & Palomba, F. (2015). An experimental investigation on the innate relationship between quality and refactoring. Journal of Systems and Software, 107, 1–14.
[13] Kannangara, S. H., & Wijayanayake, W. M. J. I. (2015). An empirical evaluation of the impact of refactoring on software quality. arXiv preprint arXiv:1502.03526.
[14] Tsantalis, N., Chaikalis, T., & Chatzigeorgiou, A. (2013). JDeodorant: Identification and removal of type-checking bad smells. Proceedings of ICSM, 329–338.
[15] AlOmar, E. A., Mkaouer, M. W., Ouni, A., & Kessentini, M. (2019). On the impact of refactoring on design quality metrics. arXiv preprint arXiv:1907.04797.
[16] Dig, D., Comertoglu, C., Marinov, D., & Johnson, R. (2009). Automated detection of refactorings in evolving components. Proceedings of ECOOP, 404–428.
[17] Bavota, G., Russo, B., & Oliveto, R. (2012). Improving API usability through refactoring. Proceedings of ICSE, 131–140.
[18] Choi, E., Bruno, E., & Cha, S. (2017). Automatic refactoring of legacy code for improving maintainability. Information and Software Technology, 83, 1–14.
[19] Alves, T. L., Ypma, C., & Visser, J. (2010). Deriving metric thresholds from benchmark data. Proceedings of ICSM, 1–10.
[20] Ouni, A., Kessentini, M., Sahraoui, H., & Hamdi, M. S. (2013). Search-based refactoring using recorded code changes. Journal of Systems and Software, 86(10), 2647–2662.
[21] Ratzinger, J., Sigmund, T., & Gall, H. C. (2008). On the relation of refactorings and software quality. Proceedings of ICSM, 35–44.
Downloads
How to Cite
[1]F. Diop, “Data-Centric Refactoring: Techniques for Improving Model Quality via Codebase Changes”, IJAIDT, vol. 2, no. 1, pp. 01–15, Apr. 2019, doi: 10.67228/30713315/IJAIDT-2019PI6J4F.