Transformer-Based Architectures for Intelligent ETL Processes

  • Authors

    • Katrina Keif Faculty of Information Technology, Brno University of Technology, Czech Republic Author

    DOI:

    https://doi.org/10.67228/30713315/IJAIDT-2022PII9A1N

    Published 12-03-2022

  • Transformer Architecture, ETL, Data Engineering, Semantic Data Integration, Intelligent Data Pipelines, Schema Inference, Data Transformation, BERT, T5, Big Data

    Issue

    Section

    Articles

    How to Cite

    [1]
    K. Keif, “Transformer-Based Architectures for Intelligent ETL Processes”, IJAIDT, vol. 5, no. 2, pp. 01–12, Dec. 2022, doi: 10.67228/30713315/IJAIDT-2022PII9A1N.
  • Abstract

    Traditional Extract, Transform, Load (ETL) processes are fundamental to data integration and management, yet they often struggle with scalability, adaptability, and semantic understanding of diverse data sources. This paper explores the application of Transformer-based architectures to design intelligent ETL pipelines capable of automating complex data transformation tasks, handling schema drift, and improving data quality through contextual semantic analysis. We propose a modular ETL framework powered by state-of-the-art Transformer models such as BERT and T5, and demonstrate how these models enhance schema inference, data mapping, and transformation logic. Comparative evaluations show significant improvements in accuracy and adaptability over traditional ETL systems. The paper also discusses challenges, limitations, and directions for future research in the convergence of AI and data engineering.

  • References

    [1] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.

    [2] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL-HLT.

    [3] Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., ... & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21(140), 1-67.

    [4] Xu, Z., Liu, H., Gong, Z., & Wang, H. (2021). Applying transformers for entity resolution in knowledge graphs. Information Sciences, 546, 241-255.

    [5] Lample, G., Conneau, A., Denoyer, L., & Ranzato, M. (2019). Cross-lingual language model pretraining. NeurIPS.

    [6] Wu, J., Lei, Y., & Wang, J. (2021). Intelligent ETL: A survey on emerging AI-driven data integration techniques. IEEE Transactions on Knowledge and Data Engineering.

    [7] Li, X., Zhou, J., & Feng, X. (2022). Schema mapping with transformers: A deep learning approach for data integration. VLDB Journal.

    [8] Doshi, R., Monga, A., & Grewal, K. (2020). Document understanding with transformers: A comprehensive review. ACM Computing Surveys.

    [9] Chen, M., Li, X., & Wong, A. (2022). Transformer-based entity resolution for heterogeneous data sources. IEEE Big Data.

    [10] Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234-1240.

    [11] Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., & Le, Q. V. (2019). XLNet: Generalized autoregressive pretraining for language understanding. NeurIPS.

    [12] Zhang, Y., & Yang, Q. (2017). A survey on multi-task learning. IEEE Transactions on Knowledge and Data Engineering, 34(12), 5586-5609.

    [13] Chung, J., & Glass, J. (2020). Streaming data integration with transformers for real-time analytics. IEEE ICDM.

    [14] Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., & Fidler, S. (2015). Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. ICCV.

    [15] Rajpurkar, P., Jia, R., & Liang, P. (2018). Know what you don't know: Unanswerable questions for SQuAD. ACL.

  • Downloads