AI-Driven Data Migration Strategies for Enterprise System

  • Authors

    • Dr. Matteo Rossi Professor, Sapienza University of Rome, Italy. Author
    • Jennifer Clark Senior Data Scientist, Innovate Analytics, USA. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2019PI1X7K

    Published 01-15-2019

  • Artificial Intelligence, Data Migration, Enterprise Systems, Machine Learning, ETL, Data Transformation, Data Quality, Automation

    Issue

    Section

    Articles

    How to Cite

    [1]
    M. Rossi and J. Clark, “AI-Driven Data Migration Strategies for Enterprise System”, IJDEIC, vol. 2, no. 1, pp. 01–12, Jan. 2019, doi: 10.67228/30715717/IJDEIC-2019PI1X7K.
  • Abstract

    In modernization of an enterprise system, data migration is a very important aspect, which involves the movement of data between storage systems, formats or applications. Conventional migration methods tend to be both labor-intensive, full of errors, and unable to scale to large amounts of heterogeneous data. Artificial Intelligence (AI) has become a revolutionary technology that could make the process of data migration more efficient, precise, and flexible. The given paper is a detailed discussion of AI-based data migration plans, with attention to the methods that have been developed. It discusses how machine learning, natural language processing and intelligent automation can be integrated into solving problems like schema mapping, data cleansing, transformation and validation. The paper will discuss the shortcomings of traditional Extract, Transform, Load (ETL) systems and emphasize how AI-driven solutions can streamline migration processes by predicting migration success and making decisions automatically. Moreover, the paper compares different AI methods, such as supervised and unsupervised learning, systems controlled by rules, and heuristic optimization, in enhancing the quality of data and minimizing migration risks. A systematic approach is suggested to apply AI-based migration, which includes data profiling, model training, iterative validation, and feedback. The findings reveal that AI-based migration strategies are very effective in enhancing accuracy of the migration process, lessening downtime, and minimizing human intervention. Per cent-based evaluation measures are the metrics of comparative analysis where the effectiveness of AI approaches will be demonstrated in comparison to traditional ones. The paper wraps up by highlighting future directions, such as incorporating deep learning and autonomous migration systems, and the need to align AI strategies with enterprise governance and compliance needs.

  • References

    [1] Rahm, E., & Do, H. H. (2000). Data cleaning: Problems and current approaches. IEEE Data Engineering Bulletin, 23(4), 3–13.

    [2] Doan, A., Halevy, A., & Ives, Z. (2012). Principles of Data Integration. Morgan Kaufmann.

    [3] Kimball, R., & Caserta, J. (2004). The Data Warehouse ETL Toolkit. Wiley Publishing.

    [4] Bernstein, P. A., & Rahm, E. (2001). Data warehouse scenarios for model management. Proceedings of ER Conference.

    [5] Batini, C., & Scannapieco, M. (2006). Data Quality: Concepts, Methodologies and Techniques. Springer.

    [6] Doan, A., Domingos, P., & Halevy, A. (2003). Learning to match the schemas of data sources: A multistrategy approach. Machine Learning Journal, 50(3), 279–301.

    [7] Noy, N. F. (2004). Semantic integration: A survey of ontology-based approaches. ACM SIGMOD Record, 33(4), 65–70.

    [8] Shvaiko, P., & Euzenat, J. (2013). Ontology matching: State of the art and future challenges. IEEE Transactions on Knowledge and Data Engineering, 25(1), 158–176.

    [9] Stonebraker, M., et al. (2018). Data curation at scale: The data wrangling challenge. CIDR Conference.

    [10] Chu, X., Ilyas, I. F., Krishnan, S., & Wang, J. (2016). Data cleaning: Overview and emerging challenges. ACM SIGMOD Record, 45(3), 19–24.

    [11] Kandel, S., Paepcke, A., Hellerstein, J. M., & Heer, J. (2011). Wrangler: Interactive visual specification of data transformation scripts. CHI Conference.

    [12] Rekatsinas, T., Chu, X., Ilyas, I. F., & Ré, C. (2017). HoloClean: Holistic data repairs with probabilistic inference. VLDB Conference.

    [13] He, Y., & Garcia-Molina, H. (2008). Learning from data to resolve entity resolution. ACM KDD Conference.

    [14] Dong, X. L., & Srivastava, D. (2015). Big data integration. Morgan & Claypool Publishers.

    [15] Abedjan, Z., Golab, L., & Naumann, F. (2016). Profiling relational data: A survey. VLDB Journal, 24(4), 557–581.

  • Downloads