A Review of Modern Database Systems for Data Science Applications
-
DOI:
https://doi.org/10.67228/30713498/IJADSMC-2023PII0Q6VPublished 12-04-2023
Database Systems, Data Science, NoSQL, NewSQL, Big Data Analytics, Distributed Databases, Cloud Computing Issue
Section
ArticlesHow to Cite
[1]A. B. Ali, “A Review of Modern Database Systems for Data Science Applications”, IJADSMC, vol. 6, no. 2, pp. 01–14, Dec. 2023, doi: 10.67228/30713498/IJADSMC-2023PII0Q6V.Abstract
Modern data mining applications require scalable and high-performance database systems capable of handling structured, semi-structured, and unstructured data. Traditional RDBMS, designed for transactional consistency, face limitations in scalability and performance for big data analytics. As a result, modern systems such as NoSQL, NewSQL, distributed file systems, and cloud-native platforms have emerged, offering features like horizontal scalability, schema flexibility, and real-time processing. This paper provides a comprehensive overview and comparison of these database systems, focusing on their architecture, data models, and suitability for analytical workloads. It also examines their role in data science pipelines, including data ingestion, preprocessing, model training, and deployment. Key trade-offs based on the CAP theorem are discussed, along with performance metrics such as scalability, latency, and fault tolerance. The study highlights that no single database solution fits all scenarios, and hybrid architectures with polyglot persistence are increasingly adopted. The paper concludes by identifying research gaps in database optimization for machine learning, data governance, and real-time analytics, offering guidance for selecting appropriate systems in data-driven environments.
References
[1] Codd, E. F. (1970). A relational model of data for large shared data banks. Communications of the ACM, 13(6), 377–387.
[2] Date, C. J. (2004). An Introduction to Database Systems (8th ed.). Pearson Education.
[3] Stonebraker, M., & Hellerstein, J. M. (2005). What goes around comes around. In Readings in Database Systems (4th ed.). MIT Press.
[4] Brewer, E. A. (2000). Towards robust distributed systems. Proceedings of the 19th Annual ACM Symposium on Principles of Distributed Computing (PODC).
[5] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113.
[6] Chang, F., et al. (2008). Bigtable: A distributed storage system for structured data. ACM Transactions on Computer Systems, 26(2), 1–26.
[7] Amazon Web Services. (2007). Dynamo: Amazon’s highly available key-value store. Proceedings of SOSP.
[8] Sadalage, P. J., & Fowler, M. (2013). NoSQL Distilled: A Brief Guide to the Emerging World of Polyglot Persistence. Addison-Wesley.
[9] Pokorny, J. (2013). NoSQL databases: A step to database scalability in web environment. International Journal of Web Information Systems, 9(1), 69–82.
[10] Stonebraker, M., et al. (2010). The end of an architectural era (It’s time for a complete rewrite). Proceedings of VLDB, 2(2), 1150–1160.
[11] Pavlo, A., et al. (2017). What’s really new with NewSQL? ACM SIGMOD Record, 45(2), 45–55.
[12] Abadi, D. J., et al. (2013). The design and implementation of modern column-oriented database systems. Foundations and Trends in Databases, 5(3), 197–280.
[13] Melnik, S., et al. (2010). Dremel: Interactive analysis of web-scale datasets. Proceedings of VLDB, 3(1), 330–339.
[14] Kraska, T., et al. (2018). The case for learned index structures. Proceedings of SIGMOD, 489–504.
Downloads
How to Cite
[1]A. B. Ali, “A Review of Modern Database Systems for Data Science Applications”, IJADSMC, vol. 6, no. 2, pp. 01–14, Dec. 2023, doi: 10.67228/30713498/IJADSMC-2023PII0Q6V.