Self-built K8s clusters have no shortage of pitfalls, especially as the Node count keeps growing — problems gradually surfaced. This blog mainly introduces two problems we met after using flannel and their solutions. The problems weren’t actually serious; they just involved the underlying structure, so changes must be made carefully.
Problem 1: flannel’s OOM Problem
The Official Configuration
The image below is the official configuration; you can see the default resource setting gives only 50M of memory:

1 | kubectl -n kube-system describe ds kube-flannel-ds-amd64 |
The Problem We Met
When our machine count exceeded 100, flannel kept dying with OOM…
1 | Feb 9 04:52:44 kernel: [37630249.323630] Memory cgroup out of memory: Kill process 33838 (flanneld) score 1653 or sacrifice child |
The data collected via Prometheus also shows the container’s memory usage was not optimistic:

There’s no good solution — the only option was adjusting the resource limits.
Problem 2: flannel’s NIC Designation Problem
Problem Background
Because the machines we use are fairly mixed, their NICs also differ; we hit the problem below when starting to build the cluster.
1 | > our virtual machines' NICs only have internal addresses starting with `10` |
The problem this brings is flannel communication: with multiple NICs and none specified at startup, flannel picks a default NIC. For virtual machines this doesn’t matter, but for physical machines flannel finds eth0 — the external NIC. flannel sending data through the wrong NIC — captured packets show flannel using the public NIC to send internal data, which gets dropped by the switch. I won’t paste the specific images; the IPs are company confidential.
The concrete fix is ensuring flannel uses the correct NIC, which requires specifying the --iface and --iface-regex parameters at startup:
we have few virtual machines and many physical ones. Besides eth1 there are also NIC names like bond1, so for virtual machines we uniformly renamed their eth0 to eth1,
then specified the configuration -iface-regex=eth1|bond1 — friendlier for adding physical machines later.
The problem seems to end here, but as flannel frequently OOM-restarted, our configuration problem was exposed.
We Found flannel Couldn’t Restart Normally After OOM
1 | NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES |
Why didn’t it appear at first, yet happens on restart? The problem lies in the regular expression. When K8s starts containers on a machine, it creates virtual NICs. These NICs’ names look like veth17f90f70@if3 — such NIC names also match the regular expression, preventing flannel from starting. The temporary solution is moving the containers off the machine;
the vethxxx NICs are automatically deleted, and flannel automatically recovers.
Of course the fundamental solution is modifying the regex configuration: - -iface-regex="^(bond1|eth1)$", making flannel match NIC names more precisely.
flannel Configuration Update and Verification
Update Preparation
Not quite knowing whether flannel handles traffic, I was a bit afraid when updating flannel — until I saw the architecture here.

flannel’s function is mainly modifying the routing tables on machines — that is, as long as machines aren’t added or removed, it doesn’t matter if flannel dies, because the routing tables need no modification.
Update
We have over 100 nodes; the whole cluster update process lasted about 1+ hour, and services were completely normal during the update.
Verifying Availability
Memory usage:

To verify flannel was usable, we deleted one node and observed the routing tables on other machines being modified in sync.
Summary
- Problems appearing isn’t scary; what matters is good monitoring with timely alarms. Our kube-system monitoring hadn’t been very good; I found the flannel-keeps-failing-to-start problem while checking.
- Before using yaml files provided by others, pay attention to the resource settings. Prometheus has this kind of problem too — it has high memory requirements.
- If the budget is sufficient, don’t build your own cluster — there are quite a few ops problems, and if an unsolvable one appears it’s very troublesome, like in the last article: Investigating a Kubernetes machine kernel problem
I hope our experience can help readers using K8s.