Millet Porridge

English version of https://corvo.myseu.cn

0%

Two flannel Problems in Our K8s Cluster

Self-built K8s clusters have no shortage of pitfalls, especially as the Node count keeps growing — problems gradually surfaced. This blog mainly introduces two problems we met after using flannel and their solutions. The problems weren’t actually serious; they just involved the underlying structure, so changes must be made carefully.

Problem 1: flannel’s OOM Problem

The Official Configuration

The image below is the official configuration; you can see the default resource setting gives only 50M of memory:

20220218220306

1
2
3
4
5
6
7
kubectl -n kube-system describe ds kube-flannel-ds-amd64
Limits:
cpu: 100m
memory: 50Mi
Requests:
cpu: 100m
memory: 50Mi

The Problem We Met

When our machine count exceeded 100, flannel kept dying with OOM…

1
Feb  9 04:52:44  kernel: [37630249.323630] Memory cgroup out of memory: Kill process 33838 (flanneld) score 1653 or sacrifice child

The data collected via Prometheus also shows the container’s memory usage was not optimistic:

20220218220628

There’s no good solution — the only option was adjusting the resource limits.

Problem 2: flannel’s NIC Designation Problem

Problem Background

Because the machines we use are fairly mixed, their NICs also differ; we hit the problem below when starting to build the cluster.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
> our virtual machines' NICs only have internal addresses starting with `10`
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
inet 127.0.0.1/8 scope host lo
valid_lft forever preferred_lft forever
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1400 qdisc pfifo_fast state UP group default qlen 1000
link/ether 52:xx:xx:xx:77:0c brd ff:ff:ff:ff:ff:ff
inet 10.xxx.xxx.xxx/26 brd 10.xxx.xxx.xxx scope global eth0
valid_lft forever preferred_lft forever

> physical machines' NICs have both public addresses starting with `59` and internal addresses starting with `10`, and the NIC is named eth1
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
inet 127.0.0.1/8 scope host lo
valid_lft forever preferred_lft forever
2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP group default qlen 1000
link/ether 8c:xx:xx:xx:xx:xx brd ff:ff:ff:ff:ff:ff
inet 59.xxx.xxx.xxx/24 brd 59.xxx.xxx.xx scope global eth0
valid_lft forever preferred_lft forever
3: eth1: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP group default qlen 1000
link/ether 8c:xx:xx:xx:xx:38 brd ff:ff:ff:ff:ff:ff
inet 10.xxx.xxx.xxx/24 brd 10.xxx.xxx.xxx scope global eth1
valid_lft forever preferred_lft forever

The problem this brings is flannel communication: with multiple NICs and none specified at startup, flannel picks a default NIC. For virtual machines this doesn’t matter, but for physical machines flannel finds eth0 — the external NIC. flannel sending data through the wrong NIC — captured packets show flannel using the public NIC to send internal data, which gets dropped by the switch. I won’t paste the specific images; the IPs are company confidential.

The concrete fix is ensuring flannel uses the correct NIC, which requires specifying the --iface and --iface-regex parameters at startup: we have few virtual machines and many physical ones. Besides eth1 there are also NIC names like bond1, so for virtual machines we uniformly renamed their eth0 to eth1, then specified the configuration -iface-regex=eth1|bond1 — friendlier for adding physical machines later.

The problem seems to end here, but as flannel frequently OOM-restarted, our configuration problem was exposed.

We Found flannel Couldn’t Restart Normally After OOM

1
2
3
NAME                                                    READY   STATUS             RESTARTS   AGE     IP               NODE                                 NOMINATED NODE   READINESS GATES
kube-flannel-ds-amd64-54c5p 0/1 CrashLoopBackOff 1604 516d 10.xx.xx.xx xxxxx <none> <none>
kube-flannel-ds-amd64-cmczh 0/1 CrashLoopBackOff 89 388d 10.xx.xx.xx yyyyy <none> <none>

Why didn’t it appear at first, yet happens on restart? The problem lies in the regular expression. When K8s starts containers on a machine, it creates virtual NICs. These NICs’ names look like veth17f90f70@if3 — such NIC names also match the regular expression, preventing flannel from starting. The temporary solution is moving the containers off the machine; the vethxxx NICs are automatically deleted, and flannel automatically recovers.

Of course the fundamental solution is modifying the regex configuration: - -iface-regex="^(bond1|eth1)$", making flannel match NIC names more precisely.

flannel Configuration Update and Verification

Update Preparation

Not quite knowing whether flannel handles traffic, I was a bit afraid when updating flannel — until I saw the architecture here.

20220218224614

flannel’s function is mainly modifying the routing tables on machines — that is, as long as machines aren’t added or removed, it doesn’t matter if flannel dies, because the routing tables need no modification.

Update

We have over 100 nodes; the whole cluster update process lasted about 1+ hour, and services were completely normal during the update.

Verifying Availability

Memory usage:

20220218225126

To verify flannel was usable, we deleted one node and observed the routing tables on other machines being modified in sync.

Summary

  1. Problems appearing isn’t scary; what matters is good monitoring with timely alarms. Our kube-system monitoring hadn’t been very good; I found the flannel-keeps-failing-to-start problem while checking.
  2. Before using yaml files provided by others, pay attention to the resource settings. Prometheus has this kind of problem too — it has high memory requirements.
  3. If the budget is sufficient, don’t build your own cluster — there are quite a few ops problems, and if an unsolvable one appears it’s very troublesome, like in the last article: Investigating a Kubernetes machine kernel problem

I hope our experience can help readers using K8s.