On Adaptive Metric-Driven Load Balancing for Specialized Clusters: A Case on Using NGINX

Publication Date

1-1-2026

Document Type

Conference Proceeding

Publication Title

Communications in Computer and Information Science

Volume

2932 CCIS

DOI

10.1007/978-3-032-22193-3_9

First Page

120

Last Page

135

Abstract

Cloud computing has become an essential part of today’s digital world. Many traditional load-balancing approaches in cloud computing lack the adaptability to dynamic workloads, and many newer load-balancing algorithms work mainly on cloud simulators. This paper proposed an adaptive, metric-driven, two-tier load-balancing system that uses NGINX and Prometheus to optimize resource allocation in specialized cloud clusters. The proposed framework is built to give great performance on Google Kubernetes Engine (GKE), but it may also be deployed in local cloud environments for security. The first tier of the proposed framework uses an NGINX-based load balancer to route incoming requests based on content type, sending traffic to hardware-optimized clusters to process requests through specialized hardware. The second tier dynamically distributes load throughout each cluster by calculating pod weights based on CPU, memory, and network usage on a regular basis. This adaptive weight distribution maximizes resource consumption and responsiveness, outperforming existing fixed-weight systems. In comparison to typical multi-component solutions, our system’s lightweight architecture minimizes configuration and management overhead while scaling fluidly to address dynamic traffic patterns. Performance evaluation shows that the proposed system improves throughput by approximately 74% and almost doubles the number of requests processed per second while being a simple, scalable, adaptable solution. We believe that the proposed solution may be a practical, effectual building block for a complex cloud systems, and would contribute significantly to the future cloud systems.

Keywords

cluster-level load balancing, Content-based routing, Google Kubernetes Engine (GKE), multi-tier cloud architecture, resource allocation

Department

Computer Science

Share

COinS