Large Language Model-Assisted Metadata Engineering for Enterprise Data Platforms

  • Authors

    • David Wheeler Professor, University of Cambridge, United Kingdom. Author
    • Michael Gordon Professor of Computer Science, University of Cambridge, United Kingdom. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2025PII5X8V

    Published 12-05-2025

  • Large Language Models (LLMs), Metadata Engineering, Enterprise Data Platforms, Data Governance, Knowledge Graphs, Semantic Metadata, Retrieval-Augmented Generation (RAG), Enterprise Data Catalog, Artificial Intelligence, Data Lineage

    Issue

    Section

    Articles

    How to Cite

    [1]
    D. Wheeler and M. Gordon, “Large Language Model-Assisted Metadata Engineering for Enterprise Data Platforms”, IJDEIC, vol. 8, no. 2, pp. 01–17, Dec. 2025, doi: 10.67228/30715717/IJDEIC-2025PII5X8V.
  • Abstract

    Enterprise metadata is essential for data discovery, governance, integration, and analytics. Traditional metadata engineering relies on manual, rule-based approaches that struggle with dynamic, heterogeneous enterprise data across cloud, IoT, ERP, CRM, and data lake environments. This study proposes a Large Language Model-Assisted Metadata Engineering Framework (LLM-MEF) that automates metadata extraction, semantic enrichment, schema recommendation, lineage discovery, and governance validation using transformer-based LLMs, Retrieval-Augmented Generation (RAG), vector databases, and knowledge graphs. The framework improves metadata quality, semantic consistency, discoverability, governance compliance, and operational efficiency while reducing manual effort. Explainable AI and continuous feedback learning further enhance transparency and adaptive improvement. The proposed LLM-MEF provides a scalable and intelligent solution for enterprise metadata management, supporting modern data governance, AI, business intelligence, regulatory compliance, and digital transformation.

  • References

    [1] T. Brown et al., “Language Models are Few-Shot Learners,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 1877–1901, 2021.

    [2] J. Wei et al., “Emergent Abilities of Large Language Models,” Transactions on Machine Learning Research (TMLR), pp. 1–30, 2022.

    [3] S. Bubeck et al., “Sparks of Artificial General Intelligence: Early Experiments with GPT-4,” arXiv:2303.12712, 2023.

    [4] P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 9459–9474, 2021.

    [5] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” IEEE Intelligent Systems, vol. 36, no. 2, pp. 28–35, 2021.

    [6] H. Wang, Y. Zhang, and J. Li, “Artificial Intelligence for Enterprise Metadata Management: A Survey,” IEEE Access, vol. 10, pp. 78215–78238, 2022.

    [7] X. Chen, L. Zhao, and Y. Liu, “Knowledge Graph-Based Enterprise Data Management: Recent Advances and Challenges,” IEEE Access, vol. 11, pp. 21567–21589, 2023.

    [8] Y. Guo, P. Wang, and S. Li, “Semantic Metadata Engineering for Enterprise Data Integration Using Deep Learning,” Future Generation Computer Systems, vol. 143, pp. 184–197, 2023.

    [9] A. Hogan et al., “Knowledge Graphs,” ACM Computing Surveys, vol. 54, no. 4, pp. 1–37, 2022.

    [10] B. Min, H. Ross, E. Sulem, A. Veyseh, T. H. Nguyen, and D. Roth, “Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey,” ACM Computing Surveys, vol. 56, no. 2, pp. 1–40, 2024.

    [11] Y. Zhao, H. Chen, and X. Wu, “Explainable Artificial Intelligence for Enterprise Decision Support Systems: A Comprehensive Review,” IEEE Access, vol. 11, pp. 102531–102558, 2023.

    [12] Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, and J. Fung, “Survey of Hallucination in Natural Language Generation,” ACM Computing Surveys, vol. 55, no. 12, pp. 1–38, 2023.

    [13] S. Pan, M. Xiao, and J. Wu, “Ontology-Driven Metadata Management for Intelligent Enterprise Information Systems,” Information Systems, vol. 118, Art. no. 102281, 2024.

    [14] M. Bommasani et al., “On the Opportunities and Risks of Foundation Models,” ACM Computing Surveys, vol. 57, no. 1, pp. 1–76, 2025.

    [15] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.

    [16] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845.

  • Downloads