Automated Data Quality Assessment Using Machine Learning

  • Authors

    • Dr. Nguyen Thi Lan Department of Information Technology, Vietnam National University, Hanoi, Vietnam Author

    DOI:

    https://doi.org/10.67228/30713498/IJADSMC-2020PII4N1L

    Published 10-05-2020

  • Data Quality, Machine Learning, Data Cleaning, Anomaly Detection, Data Preprocessing, Predictive Modeling, Data Mining, Big Data Analytics

    Issue

    Section

    Articles

    How to Cite

    [1]
    N. T. Lan, “Automated Data Quality Assessment Using Machine Learning”, IJADSMC, vol. 3, no. 2, pp. 01–13, Oct. 2020, doi: 10.67228/30713498/IJADSMC-2020PII4N1L.
  • Abstract

    Data quality is a critical factor in modern data-driven systems, influencing decision-making, predictive modeling, and operational efficiency across domains such as healthcare, finance, and smart cities. Traditional data quality assessment methods, which rely on rule-based frameworks and manual reviews, are often time-consuming, error-prone, and unable to handle large and dynamic datasets. This paper proposes an automated data quality measurement framework using machine learning techniques, focusing on scalability, adaptability, and accuracy. The proposed system integrates supervised and unsupervised learning models to detect anomalies, inconsistencies, missing values, and semantic errors in structured data. It evaluates key data quality dimensions, including completeness, consistency, accuracy, timeliness, and validity, using algorithms such as Decision Trees, Random Forests, Support Vector Machines, and clustering techniques like K-Means and DBSCAN. Feature engineering is applied to capture complex data patterns, while a hybrid approach combining statistical profiling and predictive modeling enhances detection capabilities. The framework includes data preprocessing, feature extraction, model training, evaluation, and deployment stages. Experimental results using benchmark datasets demonstrate improved performance based on precision, recall, F1-score, and accuracy, outperforming traditional rule-based methods. The study highlights challenges such as data heterogeneity, scalability, and model interpretability, and suggests future improvements. Overall, the proposed approach offers a scalable, domain-independent solution that enhances data reliability and reduces manual effort, contributing to efficient data quality management.

  • References

    [1] Batini, C., Scannapieco, M. (2016). Data Quality: Concepts, Methodologies and Techniques. Springer.

    [2] Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5–33.

    [3] Redman, T. C. (1998). The impact of poor data quality on the typical enterprise. Communications of the ACM, 41(2), 79–82.

    [4] Rahm, E., & Do, H. H. (2000). Data cleaning: Problems and current approaches. IEEE Data Engineering Bulletin, 23(4), 3–13.

    [5] Kandel, S., Paepcke, A., Hellerstein, J. M., & Heer, J. (2011). Wrangler: Interactive visual specification of data transformation scripts. CHI Conference Proceedings.

    [6] Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), 1–58.

    [7] Hawkins, D. M. (1980). Identification of Outliers. Springer.

    [8] Aggarwal, C. C. (2017). Outlier Analysis (2nd ed.). Springer.

    [9] Breunig, M. M., Kriegel, H. P., Ng, R. T., & Sander, J. (2000). LOF: Identifying density-based local outliers. ACM SIGMOD Conference.

    [10] Quinlan, J. R. (1986). Induction of decision trees. Machine Learning, 1(1), 81–106.

    [11] Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297.

    [12] Jain, A. K. (2010). Data clustering: 50 years beyond K-means. Pattern Recognition Letters, 31(8), 651–666.

    [13] Dietterich, T. G. (2000). Ensemble methods in machine learning. International Workshop on Multiple Classifier Systems.

    [14] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 1–37.

    [15] Sadiq, S., & Indulska, M. (2017). Open data: Quality over quantity. International Journal of Information Management, 37(3), 150–154.

  • Downloads