AI-Enabled Data Profiling Techniques for Quality Improvement

  • Authors

    • Dr. Rebecca Green Associate Professor, Arizona State University, USA. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2018PI1T5Q

    Published 01-07-2018

  • Artificial Intelligence, Data Profiling, Data Quality, Machine Learning, Anomaly Detection, Data Cleaning, Big Data, Predictive Analytics, Data Governance, Deep Learning

    Issue

    Section

    Articles

    How to Cite

    [1]
    R. Green, “AI-Enabled Data Profiling Techniques for Quality Improvement”, IJDEIC, vol. 1, no. 1, pp. 01–14, Jan. 2018, doi: 10.67228/30715717/IJDEIC-2018PI1T5Q.
  • Abstract

    AI is transforming data profiling by making data quality assessment more automated, scalable, and intelligent. Traditional methods struggle with modern, complex, and diverse datasets, while AI techniques—such as machine learning, deep learning, and statistical models—can detect anomalies, recognize patterns, and predict data quality issues across both structured and unstructured data. The study reviews the evolution from rule-based profiling to adaptive AI-driven approaches, highlighting methods like clustering, classification, and natural language processing. It also proposes a framework integrating AI into key stages of data profiling, including data ingestion, preprocessing, feature extraction, anomaly detection, and quality scoring, with continuous learning and dynamic rule generation. Results show that AI-based profiling significantly improves data completeness, consistency, and timeliness compared to traditional methods. Despite challenges in implementation, AI-powered data profiling represents a major advancement in data quality management, with future directions focusing on explainable AI, real-time processing, and big data integration.

  • References

    [1] H. Müller and J. Freytag, “Problems, methods, and challenges in comprehensive data cleansing,” Prof. VLDB Endowment, vol. 3, no. 1–2, pp. 155–158, 2010.

    [2] E. Rahm and H. H. Do, “Data cleaning: Problems and current approaches,” IEEE Data Engineering Bulletin, vol. 23, no. 4, pp. 3–13, 2000.

    [3] T. Dasu and T. Johnson, Exploratory Data Mining and Data Cleaning, Wiley, 2003.

    [4] F. Naumann, “Data profiling revisited,” ACM SIGMOD Record, vol. 42, no. 4, pp. 40–49, 2014.

    [5] Z. Abedjan, L. Golab, and F. Naumann, “Profiling relational data: A survey,” VLDB Journal, vol. 24, no. 4, pp. 557–581, 2015.

    [6] P. Christen, Data Matching: Concepts and Techniques for Record Linkage, Springer, 2012.

    [7] C. Batini and M. Scannapieco, Data Quality: Concepts, Methodologies and Techniques, Springer, 2006.

    [8] J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 3rd ed., Morgan Kaufmann, 2012.

    [9] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning, MIT Press, 2016.

    [10] K. Murphy, Machine Learning: A Probabilistic Perspective, MIT Press, 2012.

    [11] T. M. Mitchell, Machine Learning, McGraw-Hill, 1997.

    [12] M. Stonebraker et al., “Data curation at scale: The data tamer system,” CIDR Conference, 2013.

    [13] A. Halevy, A. Rajaraman, and J. Ordille, “Data integration: The teenage years,” VLDB, 2006.

    [14] X. Chu, I. F. Ilyas, and P. Papotti, “Holistic data cleaning: Putting violations into context,” IEEE ICDE, 2013.

    [15] S. Krishnan et al., “ActiveClean: Interactive data cleaning for statistical modeling,” VLDB, vol. 9, no. 12, pp. 948–959, 2016.

  • Downloads