Automating Data Cleansing and Transformation in Large-Scale Data Engineering Systems
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2018PIIPQRPublished 11-03-2018
Data Engineering, Automation, Data Cleansing, Data Transformation, Machine Learning, Artificial Intelligence, Big Data, Data Pipelines Issue
Section
ArticlesHow to Cite
[1]L. Languish, “Automating Data Cleansing and Transformation in Large-Scale Data Engineering Systems”, IJAIDT, vol. 1, no. 2, pp. 01–09, Nov. 2018, doi: 10.67228/30713315/IJAIDT-2018PIIPQR.Abstract
Data engineering is fundamental in modern data-driven industries, where raw data must be processed, cleaned, and transformed before being used for analytics and decision-making. Automating data cleansing and transformation in large-scale data engineering systems has become a critical area of research and practice due to the increasing volume, velocity, and variety of data. Manual data cleansing is inefficient, prone to errors, and cannot scale to meet modern data demands. This study explores automation techniques for data cleansing and transformation using artificial intelligence (AI), machine learning (ML), and rule-based frameworks. We analyze existing methodologies and tools, discussing their effectiveness, challenges, and implementation strategies. The study focuses on the integration of automated processes within large-scale data pipelines, ensuring consistency, reliability, and accuracy. The methodology section details various approaches, including AI-driven data profiling, rule-based anomaly detection, and real-time transformation mechanisms. Results indicate that automated approaches significantly reduce processing time, minimize errors, and enhance data quality. The study also highlights challenges such as handling unstructured data, scalability issues, and integration complexities. Finally, future directions for research and improvements in automated data engineering workflows are discussed.
References
[1] Batini, C., Scannapieco, M. (2016). Data and Schema Integration Quality. Springer.
[2] Rahm, E., & Do, H. H. (2000). Data Cleaning: Problems and Current Approaches. IEEE Data Engineering Bulletin, 23(4), 3–13.
[3] Vassiliadis, P. (2009). A Survey of Extract–Transform–Load Technology. International Journal of Data Warehouse and Mining, 5(3), 1–27.
[4] Dasu, T., & Johnson, T. (2003). Exploratory Data Mining and Data Cleaning. Wiley.
[5] Kimball, R., & Caserta, J. (2011). The Data Warehouse ETL Toolkit. Wiley.
[6] Rahm, E., & Do, H. H. (2000). Data cleaning: Problems and current approaches. IEEE Data Engineering Bulletin, 23(4), 3–13.
[7] Kandel, S., Paepcke, A., Hellerstein, J. M., & Heer, J. (2011). Wrangler: Interactive visual specification of data transformation scripts. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 3363–3372).
[8] Kandel, S., Heer, J., Plaisant, C., Kennedy, J., Van Ham, F., Riche, N. H., Weaver, C., Lee, B., Brodbeck, D., & Buono, P. (2012). Research directions in data wrangling: Visualizations and transformations for usable and credible data. Information Visualization, 10(4), 271–288.
[9] Elmagarmid, A. K., Ipeirotis, P. G., & Verykios, V. S. (2007). Duplicate record detection: A survey. IEEE Transactions on Knowledge and Data Engineering, 19(1), 1–16.
[10] Deshmukh, R. R., & Wangikar, V. (2011). Data cleaning: Current approaches and issues. In Proceedings of the IEEE International Conference on Knowledge Engineering.
Downloads
How to Cite
[1]L. Languish, “Automating Data Cleansing and Transformation in Large-Scale Data Engineering Systems”, IJAIDT, vol. 1, no. 2, pp. 01–09, Nov. 2018, doi: 10.67228/30713315/IJAIDT-2018PIIPQR.