Millet Porridge

English version of https://corvo.myseu.cn

0%

A Scheme for Fine-Grained Control of Pod Deployment in Kubernetes

Problem Background

Not all Kubernetes clusters have large numbers of machines, and one Pod may occupy tens of GB of memory. I hope readers understand this reality before reading.

In a cluster I manage, one cluster has relatively few machines (about 7~8 compute nodes) and average specs (this article explains using a 64-core, 128G configuration; better or worse machines may appear in future). The cluster’s owner wants his machine resources effectively utilized, ideally around 80%, so as administrator I need to consider this problem.

Using the Default Deployment Strategy

Several applications in this cluster have very high memory usage; each Pod’s memory gradually rises after startup, with an acceptable range around 20G. For such large applications, resource requests can be neither too high nor too low, so the program’s resource configuration is as follows, with requests equal to limits.

1
2
3
4
5
6
requests:
cpu: "10000m"
memory: "21474836480"
limits:
cpu: "10000m"
memory: "21474836480"

Take a 128G node as an example: we expect each node to deploy 4 or 5 Pods, for the following reasons:

  1. Deploying 6 would push memory usage over 90% and monitoring would alarm;
  2. If all nodes deploy 5, alarms may occur during each rolling update;

The fairly ideal scheme is some nodes with 4 Pods and some with 5; during rolling updates, start from the 4-Pod nodes and replace Pods gradually.

But this brings a problem: Kubernetes’ default node selection strategy is fairly free — if a machine has resources, it has some chance of being selected for deployment.

5 Pods requesting 100G mem total leaves open this possibility: one node ends up with 6 Pods, another with 3. The default allocation strategy allows this.

We encountered such deployment results many times with the default strategy, and could only fix them manually with kubectl delete pod. In a colleague’s words: painfully tedious.

A Fairly Simple Control Strategy

In kubernetes, a node’s allocatable resources can be defined. We limit nodes to reserve 10% of resources; the kubelet parameters generated by ansible can add this:

1
--system-reserved=memory={{(ansible_memtotal_mb * 0.1) | int}}Mi

Then alarms are guaranteed not to occur, but loads across machines may still be unbalanced — this only partially solves the problem.

Fine Control of Pod Distribution

Because we deploy more than one application, and some applications need special treatment, we certainly can’t rely entirely on automatic allocation strategies. With few machines and a desire for high utilization, supporting users’ manual adjustment of Pod counts is necessary.

Regarding fine control of Pod counts per node, we researched several schemes:

  1. Pod Topology Spread Constraints

This scheme is fairly complex to implement. It introduces the concept of domains, grouping nodes — each group is a domain — and limits the Pod count deployed in each domain, e.g. Pod counts between two domains cannot differ by more than 1. Using this scheme to solve load imbalance introduces new problems: if we add new machines with better performance configurations, Pod counts still can’t differ by more than 1, so the new machines’ performance can’t be fully utilized.

Honestly, I can’t think of a scenario for this scheme. If anyone has suitable use cases and ideas, tell me in the comments — I’d like to learn too.

  1. Adding extended resources to nodes

I personally think this scheme is a compromise — the configuration isn’t too complex yet achieves the desired effect. The concrete implementation adds a new resource limit: When writing control strategies, use it together with cpu and mem:

  1. Users can manually modify node resource limits, and can also set them for specific applications
  2. When we get new machines with different configurations, we can modify this option to suitable values for the new machines

I believe this scheme (automatic selection + manual configuration) has basically solved our problem. One small downside: every time new machines are added, resources must be set, otherwise Pods cannot be allocated to the new nodes.

Summary

While solving the manual deployment problem we also discussed the scenarios Kubernetes suits better: having large numbers of servers; running microservices on the servers; and the cluster ideally keeping resource utilization below 80%, so that sudden traffic leaves spare time for scaling.

In this article I mentioned three handling schemes; everyone can choose according to their situation:

  1. Consider it when building the cluster — reserve resources on every node
  2. Pod topology spread constraints — I can’t think of a suitable scenario for now
  3. For clusters with fewer machines where users want fine-grained control, I still recommend using extended resources to limit.