Kubernetes-Based Infrastructure Optimization Through Reinforcement Learning for Dynamic Cloud Workloads

Authors

  • Md Akizur Rahman Faculty of Computer Science and Engineering, The University of New South Wales, Sydney, Australia Author

DOI:

https://doi.org/10.30750/

Keywords:

Kubernetes, Reinforcement Learning, Cloud Workloads, Infrastructure Optimization, Resource Allocation, Autoscaling, Cloud Computing

Abstract

Kubernetes has become a foundational platform for managing containerized applications across dynamic cloud environments,
yet efficient infrastructure optimization remains challenging because workload demand, resource utilization, service
dependencies, and application performance continuously change. Conventional Kubernetes scheduling and autoscaling
mechanisms primarily rely on predefined policies, thresholds, and reactive decisions, which may result in resource
overprovisioning, underutilization, performance degradation, or unnecessary operational costs. This research proposes a
reinforcement learning (RL)-based infrastructure optimization framework for dynamically managing Kubernetes workloads.
The proposed approach continuously observes cluster-level and workload-level states, including CPU and memory utilization,
pod density, request and limit configurations, latency, throughput, queue length, node availability, and resource costs. An
RL agent learns optimal infrastructure-management actions such as pod scaling, workload placement, resource allocation,
and node utilization adjustments through repeated interaction with the Kubernetes environment. The methodology integrates
telemetry collection, state representation, action-space design, reward engineering, training, policy evaluation, and controlled
deployment. Experimental evaluation can compare the proposed approach with conventional Kubernetes Horizontal Pod
Autoscaler and resource-based scheduling strategies using performance, utilization, scalability, and cost indicators. The
framework is designed to improve adaptive resource allocation while maintaining application-level service requirements.
The study demonstrates how reinforcement learning can support intelligent, continuous, and workload-aware Kubernetes
infrastructure optimization for modern cloud-native applications.

Downloads

Published

2025-10-30

Similar Articles

11-20 of 44

You may also start an advanced similarity search for this article.