Data Infrastructure Requirements for Leveraging Machine Learning in Generative AI Applications
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2019PI3T9MPublished 02-05-2019
Generative AI, Machine Learning, Data Infrastructure, Scalable Data Storage, Data Processing, Cloud Computing, Gans, Vaes, Data Governance, Distributed Computing, AI Ethics, Cloud Platforms, Real-Time Data Pipelines Issue
Section
ArticlesHow to Cite
[1]Z. Abdullahi, “Data Infrastructure Requirements for Leveraging Machine Learning in Generative AI Applications”, IJAIDT, vol. 2, no. 1, pp. 01–10, Feb. 2019, doi: 10.67228/30713315/IJAIDT-2019PI3T9M.Abstract
As generative AI continues to gain traction across industries, the importance of robust data infrastructure to support machine learning (ML) workflows becomes increasingly critical. This paper explores the key data infrastructure requirements for leveraging ML in generative AI applications. It covers essential components such as scalable data storage, data preprocessing, and real-time data pipelines, while also addressing the challenges of handling large, unstructured datasets. We examine the role of distributed computing and cloud platforms in supporting the computational needs of generative AI models like GANs and VAEs. Additionally, we discuss the importance of data governance, security, and compliance in ensuring the success and ethical application of generative AI. Through real-world use cases, we highlight how industries like healthcare, entertainment, and finance are benefiting from advanced data infrastructure to drive innovation in generative AI. Finally, the paper outlines best practices for building and optimizing data infrastructure to ensure that organizations can fully leverage the potential of machine learning in generative AI applications.
References
[1] Vogelsang, A., & Borg, M. (2019). Requirements engineering for machine learning: Perspectives from data scientists. arXiv preprint arXiv:1908.04674.
[2] Augenstein, S., McMahan, H. B., Ramage, D., Ramaswamy, S., Kairouz, P., Chen, M., Mathews, R., & Aguera y Arcas, B. (2019). Generative Models for Effective ML on Private, Decentralized Datasets. arXiv.
[3] Yang, Q., Liu, Y., Chen, T., & Tong, Y. (2019). Federated Machine Learning: Concept and Applications. arXiv.
[4] Hey, T., Butler, K., Jackson, S., & Thiyagalingam, J. (2019). Machine Learning and Big Scientific Data. arXiv.
[5] Kanter, J. M., Schreck, B., & Veeramachaneni, K. (2018). Machine Learning 2.0: Engineering Data-Driven AI Products. arXiv.
[6] Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press.
[7] Zaharia, M., Chen, A., Davidson, A., et al. (2018). Accelerating the Machine Learning Lifecycle with MLflow. IEEE Data Engineering Bulletin.
[8] Armbrust, M., Xin, R. S., Lian, C., et al. (2018). Apache Spark: A Unified Analytics Engine for Big Data Processing. Communications of the ACM.
[9] Abadi, M., Agarwal, A., Barham, P., et al. (2016). TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. arXiv.
[10] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified Data Processing on Large Clusters. Communications of the ACM.
[11] Shvachko, K., Kuang, H., Radia, S., & Chansler, R. (2010). The Hadoop Distributed File System. IEEE Symposium on Mass Storage Systems.
[12] Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. ACM SIGKDD.
[13] Radford, A., Metz, L., & Chintala, S. (2016). Unsupervised Representation Learning with Deep Convolutional GANs. arXiv.
[14] Kingma, D. P., & Welling, M. (2014). Auto-Encoding Variational Bayes. arXiv.
[15] Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). Mastering the Game of Go Without Human Knowledge. Nature.
[16] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep Learning. Nature.
Downloads
How to Cite
[1]Z. Abdullahi, “Data Infrastructure Requirements for Leveraging Machine Learning in Generative AI Applications”, IJAIDT, vol. 2, no. 1, pp. 01–10, Feb. 2019, doi: 10.67228/30713315/IJAIDT-2019PI3T9M.