Cognitive Data Engineering Frameworks for Automated Data Management
-
DOI:
https://doi.org/10.67228/30715717/IJDEIC-2018PI6W9PPublished 05-09-2018
Cognitive Data Engineering, Automated Data Management, Machine Learning, Data Pipelines, ETL Automation, Knowledge Graphs, Data Governance, Artificial Intelligence Issue
Section
ArticlesHow to Cite
[1]F. Noor and S. B. Reddy, “Cognitive Data Engineering Frameworks for Automated Data Management”, IJDEIC, vol. 1, no. 1, pp. 01–13, May 2018, doi: 10.67228/30715717/IJDEIC-2018PI6W9P.Abstract
Cognitive Data Engineering (CDE) is an advanced paradigm that integrates artificial intelligence, machine learning, and knowledge-based systems into traditional data engineering to enable automated and intelligent data management. This paper presents a Cognitive Data Engineering Framework (CDEF) designed to automate key data lifecycle processes such as ingestion, transformation, integration, quality assurance, and governance. Unlike conventional rule-based pipelines, the proposed framework adapts dynamically to data changes, anomalies, and schema evolution through self-learning and context-aware capabilities. The framework employs metadata-driven intelligence, semantic modeling, reinforcement learning, and cognitive agents within a layered architecture comprising perception, reasoning, learning, and execution. It also leverages knowledge graphs and ontologies to enhance semantic interoperability and data discovery. Experimental results demonstrate improved performance, reduced errors, and increased flexibility compared to traditional systems. Overall, the study highlights the potential of CDEFs in enabling efficient, scalable, and autonomous data management, with future scope in edge computing, real-time analytics, and self-governing data ecosystems.
References
[1] Dasu, T., & Johnson, T. (2003). Exploratory data mining and data cleaning. John Wiley & Sons.
[2] Kimball, R., & Caserta, J. (2004). The data warehouse ETL toolkit: Practical techniques for extracting, cleaning, conforming, and delivering data. Wiley.
[3] Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., & Zheng, X. (2016). TensorFlow: A system for large-scale machine learning. Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI '16), 265–283. https://doi.org/10.5555/3026877.3026899
[4] Paulheim, H. (2017). Knowledge graph refinement: A survey of approaches and evaluation methods. Semantic Web, 8(3), 489–508. https://doi.org/10.3233/SW-160218
[5] Nickel, M., Murphy, K., Tresp, V., & Gabrilovich, E. (2016). A review of relational machine learning for knowledge graphs. Proceedings of the IEEE, 104(1), 11–33. https://doi.org/10.1109/JPROC.2015.2483592
[6] Dong, X. L., Gabrilovich, E., Heitz, G., Horn, W., Murphy, K., Sun, S., Zhang, W., & Zhu, F. (2014). Knowledge Vault: A web-scale approach to probabilistic knowledge fusion. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 601–610. https://doi.org/10.1145/2623330.2623623
[7] Etzioni, O., Banko, M., Soderland, S., & Weld, D. S. (2008). Open information extraction from the web. Communications of the ACM, 51(12), 68–74. https://doi.org/10.1145/1409360.1409378
[8] Lenzerini, M. (2002). Data integration: A theoretical perspective. Proceedings of the ACM Symposium on Principles of Database Systems, 233–246. https://doi.org/10.1145/543613.543644
[9] Batini, C., & Scannapieco, M. (2016). Data and information quality: Dimensions, principles and techniques. Springer. https://doi.org/10.1007/978-3-319-24106-7
[10] Bizer, C., Heath, T., & Berners-Lee, T. (2009). Linked data—The story so far. International Journal on Semantic Web and Information Systems, 5(3), 1–22. https://doi.org/10.4018/jswis.2009081901
[11] Robinson, I., Webber, J., & Eifrem, E. (2015). Graph databases (2nd ed.). O'Reilly Media.
[12] Kleppmann, M. (2017). Designing data-intensive applications. O'Reilly Media.
[13] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113. https://doi.org/10.1145/1327452.1327492
[14] Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2010). Spark: Cluster computing with working sets. Proceedings of HotCloud, 10, 95–102.
[15] Carbone, P., Katsifodimos, A., Ewen, S., Markl, V., Haridi, S., & Tzoumas, K. (2015). Apache Flink: Stream and batch processing in a single engine. IEEE Data Engineering Bulletin, 38(4), 28–38.
Downloads
How to Cite
[1]F. Noor and S. B. Reddy, “Cognitive Data Engineering Frameworks for Automated Data Management”, IJDEIC, vol. 1, no. 1, pp. 01–13, May 2018, doi: 10.67228/30715717/IJDEIC-2018PI6W9P.