An earlier blog introduced the scheme for fine-grained pod control: we can targetedly adjust pod counts by adding custom resources — but that’s a basic primitive. What about providing it to users as a platform feature? This blog introduces how we leverage this feature to precisely control the number of pods deployed on each machine.
A Brief Review
Our PaaS platform allows users to provide machines, which we integrate into kubernetes. For these user-owned machines, they hope the machines’ performance utilization reaches an excellent level. No matter how much we rely on the program’s allocation strategy, it’s never as good as manual adjustment — hence the platform needs to provide a feature that fixes the Pod count deployed on each machine.
In the last blog I already gave a workable scheme; the concrete implementation adds a new resource limit:
When writing control strategies, use it together with cpu and mem:
System Design and Implementation
The User Interaction Part
After researching implementation schemes, I implemented a feature that could modify custom resource configurations. But after giving it to users, the feedback was that it was extremely hard to use — modifications involved a huge workload, almost completely unusable. After discussing schemes within the team, we decided: for each application, if the user wants to specify how many Pods deploy on each machine, the most simplified configuration should look like this:
1 | Application 1 configuration: |
- Users shouldn’t care about underlying implementation details or bother with custom resource limits — just give the node name and control is achieved. This is the feature users actually want.
The Implementation Part
After adding custom resources, program deployment needs one more step: declaring each machine’s resource count. The code below declares a machine’s custom resource.
1 | def k8s_update_capacity(node, key, value=None): |
For example, if machine1 needs 3 pods configured, then k8s_update_capacity('machine1', 'NS-name-app-name-version-name', 3); when creating deployments, set the corresponding resource request to 1, and 3 Pods will deploy on machine1. Of course, you need to do the same for all designated machines.
The Details
Note that our platform supports two release types: rolling updates and version updates. Let me briefly describe the difference:
- Rolling update: no concept of versions; when updating, old Pods are gradually deleted and new Pods gradually created
- Version update: when updating, a new version’s Pods are created, then the user can choose whether to discard or launch this new version
I think I could pad out a blog on the difference between the two
In our implementation, for applications using version updates the resource label should be NS-app_name-version_id;
for rolling-update applications the resource label should be NS-app_name.
Why do rolling-update applications have their own version numbers but can’t include them in the label?
A:
When our PaaS platform was first designed, it only had the version-update feature. After combining with k8s, the rolling-update feature was added, so even applications using rolling updates still have version numbers — they’re just not shown to k8s.
This concerns our resource usage: our applications occupy many resources. For example, on one machine, running 4 Pods is already the limit; we need the rolling-update strategy to delete a few and create a few, always maintaining a maximum Pod count of 4 — otherwise resource overuse may occur (alarms or the machine dying). If such applications used labels carrying version numbers, then at the same moment one machine could have two labels:
NS-app_name-old_versionandNS-app_name-new_version. This way, when updating the application, the machine’s old and new version Pod counts could sum to more than 4. Therefore rolling-update applications’ custom resources must absolutely not carry version numbers.
These two update methods’ label names also determine that we need different strategies when updating custom resources:
For version-update applications:
When deploying: we need to add custom resource configuration to the machines to be deployed to — since it carries the version number, the new version will definitely deploy to the correct machines. When taking the old version offline after the update: the resource settings corresponding to the old version number need deleting.
For rolling-update applications: When deploying: we need to update the resource configuration of ALL machines that might be deployed to, because this update may have changed the deployment machines. For example, host1 originally deployed 3 pods, and the new deployment specifies host1 no longer deploys the application — then host1’s custom resources must be deleted before deployment; otherwise, without the version-number restriction, pods could still deploy to host1. No other handling is needed after the update ends.
The Concrete Effect
I won’t say more — a colleague’s one sentence shows the effect.

This feature has now been running stably online for 2 months; the program behaves as expected.
Problems Encountered
Scaling Up and Down
After using custom resource limits, a fairly troublesome problem is scaling. The original scaling could be done by simply modifying replicas, but now the machines’ corresponding resource counts need synchronizing.
Scaling up can be done by simply syncing the machine resource configuration and then modifying the replica count, but scaling down cannot: concretely, if you lower a machine’s custom resource and then reduce the replica count, Pod deletion won’t necessarily happen on the corresponding machine — Pods on other machines may be deleted by mistake. Only redeployment can scale down.
The specific reason requires reading the replicaset code. My next blog will introduce k8s’s scale-down strategy, and along the way how to debug control-manager code. Unknowingly, through this blog I seem able to pad out 3 more blogs~
The maxunavailable Setting
Mainly for rolling-update applications: when we have only 2 Pods and maxunavailable is 25%,
only 1/4 container unavailability can be tolerated — so neither of the original 2 Pods will be deleted, and new Pods cannot be created either:
a deadlock appears. The concrete solution: when users configure, ensure maxunavailable * replica > 1.
Summary
Implementing this fine-grained control feature, my biggest feeling is: after researching a solution, you must provide features users can actually use according to their needs, rather than a hasty implementation (which costs time AND goes unused).
Also, for different deployment schemes, the implementation strategy of such basic features needs consideration. In the implementation above I omitted the cronjob and job implementations; I hope readers will consider them yourselves when needed — it should be fairly simple.

