Cognitive Data Engineering Frameworks for Automated Data Management

  • Authors

    • Dr. Fatima Noor Associate Professor, University of Malaya, Malaysia. Author
    • Dr. Suresh Babu Reddy Professor, Osmania University, India. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2018PI6W9P

    Published 05-09-2018

  • Cognitive Data Engineering, Automated Data Management, Machine Learning, Data Pipelines, ETL Automation, Knowledge Graphs, Data Governance, Artificial Intelligence

    Issue

    Section

    Articles

    How to Cite

    [1]
    F. Noor and S. B. Reddy, “Cognitive Data Engineering Frameworks for Automated Data Management”, IJDEIC, vol. 1, no. 1, pp. 01–13, May 2018, doi: 10.67228/30715717/IJDEIC-2018PI6W9P.
  • Abstract

    Cognitive Data Engineering (CDE) is an advanced paradigm that integrates artificial intelligence, machine learning, and knowledge-based systems into traditional data engineering to enable automated and intelligent data management. This paper presents a Cognitive Data Engineering Framework (CDEF) designed to automate key data lifecycle processes such as ingestion, transformation, integration, quality assurance, and governance. Unlike conventional rule-based pipelines, the proposed framework adapts dynamically to data changes, anomalies, and schema evolution through self-learning and context-aware capabilities. The framework employs metadata-driven intelligence, semantic modeling, reinforcement learning, and cognitive agents within a layered architecture comprising perception, reasoning, learning, and execution. It also leverages knowledge graphs and ontologies to enhance semantic interoperability and data discovery. Experimental results demonstrate improved performance, reduced errors, and increased flexibility compared to traditional systems. Overall, the study highlights the potential of CDEFs in enabling efficient, scalable, and autonomous data management, with future scope in edge computing, real-time analytics, and self-governing data ecosystems.

  • References

    [1] Dasu, T., & Johnson, T. (2003). Exploratory data mining and data cleaning. John Wiley & Sons.

    [2] Kimball, R., & Caserta, J. (2004). The data warehouse ETL toolkit: Practical techniques for extracting, cleaning, conforming, and delivering data. Wiley.

    [3] Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., & Zheng, X. (2016). TensorFlow: A system for large-scale machine learning. Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI '16), 265–283. https://doi.org/10.5555/3026877.3026899

    [4] Paulheim, H. (2017). Knowledge graph refinement: A survey of approaches and evaluation methods. Semantic Web, 8(3), 489–508. https://doi.org/10.3233/SW-160218

    [5] Nickel, M., Murphy, K., Tresp, V., & Gabrilovich, E. (2016). A review of relational machine learning for knowledge graphs. Proceedings of the IEEE, 104(1), 11–33. https://doi.org/10.1109/JPROC.2015.2483592

    [6] Dong, X. L., Gabrilovich, E., Heitz, G., Horn, W., Murphy, K., Sun, S., Zhang, W., & Zhu, F. (2014). Knowledge Vault: A web-scale approach to probabilistic knowledge fusion. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 601–610. https://doi.org/10.1145/2623330.2623623

    [7] Etzioni, O., Banko, M., Soderland, S., & Weld, D. S. (2008). Open information extraction from the web. Communications of the ACM, 51(12), 68–74. https://doi.org/10.1145/1409360.1409378

    [8] Lenzerini, M. (2002). Data integration: A theoretical perspective. Proceedings of the ACM Symposium on Principles of Database Systems, 233–246. https://doi.org/10.1145/543613.543644

    [9] Batini, C., & Scannapieco, M. (2016). Data and information quality: Dimensions, principles and techniques. Springer. https://doi.org/10.1007/978-3-319-24106-7

    [10] Bizer, C., Heath, T., & Berners-Lee, T. (2009). Linked data—The story so far. International Journal on Semantic Web and Information Systems, 5(3), 1–22. https://doi.org/10.4018/jswis.2009081901

    [11] Robinson, I., Webber, J., & Eifrem, E. (2015). Graph databases (2nd ed.). O'Reilly Media.

    [12] Kleppmann, M. (2017). Designing data-intensive applications. O'Reilly Media.

    [13] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113. https://doi.org/10.1145/1327452.1327492

    [14] Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2010). Spark: Cluster computing with working sets. Proceedings of HotCloud, 10, 95–102.

    [15] Carbone, P., Katsifodimos, A., Ewen, S., Markl, V., Haridi, S., & Tzoumas, K. (2015). Apache Flink: Stream and batch processing in a single engine. IEEE Data Engineering Bulletin, 38(4), 28–38.

  • Downloads