Reinforcement Learning for Adaptive Resource Management in Cloud Software

  • Authors

    • Dr. Rajesh Kumar Sharma Professor, University of Delhi, India. Author
    • Dr. Priya Natarajan Associate Professor, University of Madras, India. Author

    DOI:

    https://doi.org/10.67228/3142788X/IJMLPA-2021PII7K4B

    Published 10-05-2021

  • Reinforcement Learning, Cloud Resource Management, Adaptive Systems, Auto-Scaling, Software Performance Engineering, Deep Reinforcement Learning, Service-Level Agreements, Cloud Computing

    Issue

    Section

    Articles

    How to Cite

    [1]
    R. K. Sharma and P. Natarajan, “Reinforcement Learning for Adaptive Resource Management in Cloud Software”, IJMLPA, vol. 4, no. 2, pp. 01–10, Oct. 2021, doi: 10.67228/3142788X/IJMLPA-2021PII7K4B.
  • Abstract

    Cloud software systems operate under highly dynamic and unpredictable workloads, requiring efficient and adaptive resource management strategies to maintain performance, reliability, and cost efficiency. Traditional rule-based and heuristic resource allocation approaches often fail to respond optimally to rapid workload fluctuations and complex system interactions. This paper proposes a reinforcement learning-based adaptive resource management framework that enables cloud systems to autonomously learn optimal resource allocation policies through continuous interaction with the environment. By modeling cloud resource management as a sequential decision-making problem, the framework leverages reinforcement learning algorithms such as Q-learning, Deep Q-Networks (DQN), and policy-gradient methods to dynamically adjust computing resources including CPU, memory, and virtual machine instances. The proposed approach aims to optimize multiple objectives such as performance, cost, and service-level agreement (SLA) compliance. Experimental evaluation using simulated and real-world cloud workloads demonstrates that reinforcement learning significantly outperforms static and reactive baseline strategies in terms of resource utilization efficiency and response time stability. The results highlight the potential of reinforcement learning to enable intelligent, self-adaptive cloud resource management systems.

  • References

    [1] R. Sutton and A. Barto, Reinforcement Learning: An Introduction, MIT Press, 2018.

    [2] I. Foster et al., “Cloud computing and grid computing 360-degree compared,” Grid Computing Environments Workshop, 2008.

    [3] L. Mao and M. Humphrey, “Auto-scaling to minimize cost and meet application deadlines in cloud workflows,” SC Companion, 2011.

    [4] H. Xu and B. Li, “Dynamic cloud pricing for revenue maximization,” IEEE Transactions on Cloud Computing, 2013.

    [5] M. Mao and M. Humphrey, “A performance study on the VM startup time in the cloud,” IEEE Cloud, 2012.

    [6] Z. Zhang et al., “Dynamic resource provisioning in cloud computing,” IEEE Transactions on Parallel and Distributed Systems, 2014.

    [7] J. Dean and L. A. Barroso, “The tail at scale,” Communications of the ACM, 2013.

    [8] Y. Mao et al., “Resource management with deep reinforcement learning,” HotNets, 2016.

    [9] T. Chen et al., “Self-adaptive resource allocation using reinforcement learning,” IEEE CLOUD, 2018.

    [10] M. Ghodsi et al., “Dominant resource fairness,” ACM SIGCOMM, 2011.

  • Downloads