AI-Based Data Cleaning Techniques for Large-Scale Datasets
-
DOI:
https://doi.org/10.67228/30713498/IJADSMC-2020PII8S2XPublished 07-04-2020
Artificial Intelligence, Data Cleaning, Large-Scale Datasets, Machine Learning, Data Quality, Anomaly Detection, Data Preprocessing, Big Data Analytics Issue
Section
ArticlesHow to Cite
[1]F. Al-Farsi, “AI-Based Data Cleaning Techniques for Large-Scale Datasets”, IJADSMC, vol. 3, no. 2, pp. 01–14, Jul. 2020, doi: 10.67228/30713498/IJADSMC-2020PII8S2X.Abstract
The rapid growth of data from sources such as social media, IoT devices, and enterprise systems has increased the need for efficient data preprocessing, particularly data cleaning. Traditional rule-based and manual methods are no longer sufficient for handling large-scale, complex datasets. This paper explores AI-based data cleaning techniques, including supervised, unsupervised, reinforcement, and deep learning approaches, for tasks such as missing value imputation, anomaly detection, and duplicate removal. The proposed framework integrates machine learning with distributed computing to improve scalability and efficiency. Experimental results on benchmark datasets show that AI-based methods outperform traditional approaches in terms of accuracy, precision, recall, and F1-score. Despite challenges like computational complexity, interpretability, and privacy concerns, AI-driven data cleaning offers a powerful solution for improving data quality in large-scale analytics and decision-making systems.
References
[1] E. Rahm and H. H. Do, “Data cleaning: Problems and current approaches,” IEEE Data Engineering Bulletin, vol. 23, no. 4, pp. 3–13, 2000.
[2] J. Widom, “Data cleaning: Problems and current approaches,” IEEE Data Engineering Bulletin, vol. 23, no. 4, pp. 3–13, 2000.
[3] M. Stonebraker and I. F. Ilyas, “Data integration: The current status and the way forward,” IEEE Data Engineering Bulletin, vol. 33, no. 3, pp. 3–9, 2010.
[4] I. F. Ilyas and X. Chu, Trends in Cleaning Relational Data: Consistency and Deduplication, Now Publishers Inc., 2015.
[5] P. Bohannon, M. Flaster, W. Fan, and R. Rastogi, “A cost-based model and effective heuristic for repairing constraints by value modification,” in Proc. ACM SIGMOD, 2005, pp. 143–154.
[6] X. Chu, I. F. Ilyas, and P. Papotti, “Holistic data cleaning: Putting violations into context,” in Proc. IEEE ICDE, 2013, pp. 458–469.
[7] J. Van den Broeck et al., “Data cleaning: Detecting, diagnosing, and editing data abnormalities,” PLOS Medicine, vol. 2, no. 10, pp. 966–970, 2005.
[8] T. Dasu and T. Johnson, Exploratory Data Mining and Data Cleaning, Wiley, 2003.
[9] C. C. Aggarwal, “Outlier analysis,” Springer, 2017.
[10] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
[11] C. Cortes and V. Vapnik, “Support-vector networks,” Machine Learning, vol. 20, no. 3, pp. 273–297, 1995.
[12] M. Ester, H. P. Kriegel, J. Sander, and X. Xu, “A density-based algorithm for discovering clusters in large spatial databases with noise,” in Proc. KDD, 1996, pp. 226–231.
[13] F. T. Liu, K. M. Ting, and Z. H. Zhou, “Isolation forest,” in Proc. IEEE ICDM, 2008, pp. 413–422.
[14] D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in Proc. ICLR, 2014.
[15] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning, MIT Press, 2016.
Downloads
How to Cite
[1]F. Al-Farsi, “AI-Based Data Cleaning Techniques for Large-Scale Datasets”, IJADSMC, vol. 3, no. 2, pp. 01–14, Jul. 2020, doi: 10.67228/30713498/IJADSMC-2020PII8S2X.