Machine Learning-Based Storage Optimization in Distributed Databases
-
DOI:
https://doi.org/10.67228/30715717/IJDEIC-2019PII1R7WPublished 08-09-2019
Distributed Databases, Machine Learning, Storage Optimization, Data Placement, Replication Strategy, Predictive Analytics, Resource Utilization Issue
Section
ArticlesHow to Cite
[1]E. Petrova, “Machine Learning-Based Storage Optimization in Distributed Databases”, IJDEIC, vol. 2, no. 2, pp. 01–12, Aug. 2019, doi: 10.67228/30715717/IJDEIC-2019PII1R7W.Abstract
The rapid growth of data from sources like social media, IoT, enterprise systems, and cloud services has increased the need for efficient distributed database systems. However, challenges such as data redundancy, load imbalance, latency, and poor resource utilization persist. This paper explores machine learning (ML)-based storage optimization techniques that improve storage allocation, replication, and data placement. Unlike traditional heuristic methods, ML enables adaptive and data-driven optimization using techniques like clustering, regression, classification, and reinforcement learning. The study categorizes these approaches into predictive data placement, intelligent replication management, and anomaly detection. Results show that ML-based methods can enhance storage efficiency by up to 35%, while improving scalability, latency, and fault tolerance. Overall, the paper highlights ML as a key enabler for intelligent, self-optimizing storage systems, while noting challenges such as scalability, training overhead, and data privacy.
References
[1] Jeffrey Dean and Sanjay Ghemawat, “MapReduce: Simplified Data Processing on Large Clusters,” OSDI, 2004.
[2] Konstantin Shvachko et al., “The Hadoop Distributed File System,” IEEE MSST, 2010.
[3] Avinash Lakshman and Prashant Malik, “Cassandra: A Decentralized Structured Storage System,” ACM SIGOPS, 2010.
[4] Giuseppe DeCandia et al., “Dynamo: Amazon’s Highly Available Key-Value Store,” SOSP, 2007.
[5] Matei Zaharia et al., “Apache Spark: A Unified Engine for Big Data Processing,” Communications of the ACM, 2016.
[6] Kai Ren et al., “Heterogeneity-Aware Resource Allocation and Scheduling in the Cloud,” IEEE Transactions on Cloud Computing, 2015.
[7] Yuxiong He et al., “Real-Time Scheduling for Data-Intensive Applications,” HPDC, 2011.
[8] Tim Kraska et al., “The Case for Learned Index Structures,” SIGMOD, 2018.
[9] Ming Zhao et al., “MATRIX: Adaptive Data Placement in Distributed Systems,” IEEE INFOCOM, 2015.
[10] Song Jiang et al., “Improving Data Locality in Distributed Systems,” Cluster Computing, 2012.
[11] Xiaoyong Du et al., “Prediction-Based Data Placement Using Time Series Models,” Journal of Systems Architecture, 2016.
[12] Richard Sutton and Andrew Barto, “Reinforcement Learning: An Introduction,” MIT Press, 1998.
[13] Hado van Hasselt et al., “Deep Reinforcement Learning for Dynamic Resource Allocation,” AAAI, 2016.
[14] Zhenhua Guo et al., “Dynamic Data Replication Strategies in Distributed Storage Systems,” IEEE Transactions on Parallel and Distributed Systems, 2014.
[15] Abadi Daniel et al., “The Design and Implementation of Modern Column-Oriented Database Systems,” Foundations and Trends in Databases, 2013.
Downloads
How to Cite
[1]E. Petrova, “Machine Learning-Based Storage Optimization in Distributed Databases”, IJDEIC, vol. 2, no. 2, pp. 01–12, Aug. 2019, doi: 10.67228/30715717/IJDEIC-2019PII1R7W.