Millet Porridge

English version of https://corvo.myseu.cn

0%

DNS Resolution Implementation in Kubernetes

I’ve been handling DNS resolution problems in Kubernetes recently — a good chance to learn how the DNS server in Kubernetes works. The dns server problem I handled will get its own blog later.

My understanding of the resolution process is fairly shallow; I’ll only introduce the configuration contents.

DNS Overview in a Pod

As everyone knows, DNS servers convert domain names to IPs (for why the conversion is needed, I suggest reviewing the 7-layer network model). On Linux servers, dns resolution configuration lives in /etc/resolv.conf — pods are no exception. Here’s the configuration in a certain Pod:

1
2
3
nameserver 10.96.0.10
search kube-system.svc.cluster.local svc.cluster.local cluster.local
options ndots:5

Suppose we normally want to modify our own machine’s dns server, e.g. to 8.8.8.8, we’d modify it like this:

1
2
nameserver 8.8.8.8
nameserver 8.8.4.4

To debug a DNS server and test returned results, you can use the dig tool:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
> dig baidu.com @8.8.8.8

; <<>> DiG 9.16.10 <<>> baidu.com @8.8.8.8
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 5114
;; flags: qr rd ra; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1

;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags:; udp: 512
;; QUESTION SECTION:
;baidu.com. IN A

;; ANSWER SECTION:
baidu.com. 159 IN A 39.156.69.79
baidu.com. 159 IN A 220.181.38.148

;; Query time: 10 msec
;; SERVER: 8.8.8.8#53(8.8.8.8)
;; WHEN: Tue Jan 12 09:26:13 HKT 2021
;; MSG SIZE rcvd: 70

The DNS Server – nameserver

Let’s start with nameserver 10.96.0.10 — why does requesting this address perform DNS resolution? The answer is iptables. I excerpt only udp port 53; the following can be obtained via iptables-save.

1
2
-A KUBE-SERVICES -d 10.96.0.10/32 -p udp -m comment --comment "kube-system/kube-dns:dns cluster IP" -m udp --dport 53 -j KUBE-SVC-TCOU7JCQXEZGVUNU
# briefly explained: this rule means if the destination is 10.96.0.10's udp port 53, jump to the chain `KUBE-SVC-TCOU7JCQXEZGVUNU`

Let’s look at this chain KUBE-SVC-TCOU7JCQXEZGVUNU:

1
2
3
4
5
6
7
8
-A KUBE-SVC-TCOU7JCQXEZGVUNU -m statistic --mode random --probability 0.50000000000 -j KUBE-SEP-Q3HNNZPXUAYYDXW2
-A KUBE-SVC-TCOU7JCQXEZGVUNU -j KUBE-SEP-BBR3Z5NWFGXGVHEZ

-A KUBE-SEP-Q3HNNZPXUAYYDXW2 -p udp -m udp -j DNAT --to-destination 172.32.3.219:53
-A KUBE-SEP-BBR3Z5NWFGXGVHEZ -p udp -m udp -j DNAT --to-destination 172.32.6.239:53

# connecting the earlier rule, these rules' complete meaning is:
# on this machine, traffic sent to 10.96.0.10:53 is half forwarded to 172.32.3.219:53, the other half to 172.32.6.239:53

Kubernetes’ deployment

Now look at our kubernetes pods’ IP addresses — that is, dns requests actually reach our coredns containers to be processed.

1
2
3
> kubectl -n kube-system get pods -o wide | grep dns
coredns-646bc69b8d-jd22w 1/1 Running 0 57d 172.32.6.239 m1 <none> <none>
coredns-646bc69b8d-p8pqq 1/1 Running 8 315d 172.32.3.219 m2 <none> <none>

The Concrete Implementation of Services in Kubernetes

Checking the corresponding service, you can see the iptables on the machines above are actually the concrete implementation of the service.

1
2
> kubectl -n kube-system get svc | grep dns
kube-dns ClusterIP 10.96.0.10 <none> 53/UDP,53/TCP,9153/TCP 398d

Some may wonder: now 2 pods split traffic evenly — what about 3 or 4 pods? How does iptables forward then? I happened to have this question, so I added 2 more pods to see how iptables splits traffic evenly across 4 pods.

This is the final implementation:

1
2
3
4
-A KUBE-SVC-TCOU7JCQXEZGVUNU -m statistic --mode random --probability 0.25000000000 -j KUBE-SEP-HTZHQHQPOHVVNWZS
-A KUBE-SVC-TCOU7JCQXEZGVUNU -m statistic --mode random --probability 0.33333333349 -j KUBE-SEP-3VNFB2SPYQJRRPK6
-A KUBE-SVC-TCOU7JCQXEZGVUNU -m statistic --mode random --probability 0.50000000000 -j KUBE-SEP-Q3HNNZPXUAYYDXW2
-A KUBE-SVC-TCOU7JCQXEZGVUNU -j KUBE-SEP-BBR3Z5NWFGXGVHEZ

These statements should mean:

  1. The first 1/4 of traffic goes to one chain, leaving 3/4
  2. Of the remaining 3/4, 1/3 goes to one chain, leaving 2/4
  3. Of the remaining 2/4, 1/2 goes to one chain, leaving 1/4
  4. The last 1/4 goes to one chain

Traffic is evenly split this way — quite clever. This way, 5 or 10 pods can also be divided successively.

Parsing Other resolv.conf Parameters

1
2
search kube-system.svc.cluster.local svc.cluster.local cluster.local
options ndots:5

For a detailed introduction see here: resolv.conf manual. Let me briefly state my understanding.

The search Parameter

Without this search parameter, when we look up:

1
2
> ping kube-dns
ping: kube-dns: Name or service not known

After adding the search parameter, look up again:

1
2
> ping kube-dns
PING kube-dns.kube-system.svc.psigor-dev.nease.net (10.96.0.10) 56(84) bytes of data.

You can see that when resolving a domain, if the given domain can’t be found, the suffixes after search are appended for lookup (if it ends with ., like kube-dns., such FQDN domains won’t be retried).

search‘s job is trying for us. Used in kubernetes, configuring kube-system.svc.cluster.local svc.cluster.local cluster.local means when we ping abc, queries proceed like this:

1
2
3
4
[INFO] 10.202.37.232:50940 - 51439 "A IN abc.kube-system.svc.cluster.local. udp 51 false 512" NXDOMAIN qr,aa,rd 144 0.000114128s
[INFO] 10.202.37.232:51823 - 54524 "A IN abc.svc.cluster.local. udp 39 false 512" NXDOMAIN qr,aa,rd 132 0.000124048s
[INFO] 10.202.37.232:41894 - 15434 "A IN abc.cluster.local. udp 35 false 512" NXDOMAIN qr,aa,rd 128 0.000092304s
[INFO] 10.202.37.232:40357 - 43160 "A IN abc. udp 21 false 512" NOERROR qr,aa,rd,ra 94 0.000163406s

ndots and Its Optimization Problem

The search configuration must be used together with ndots. The default ndots is 1; its function: if the queried domain contains fewer dots than this value, suffixes from the search domains are tried first.

1
2
3
4
Resolver queries having fewer than
ndots dots (default is 1) in them will be attempted using
each component of the search path in turn until a match is
found.

A Concrete Example

Suppose our dns configuration is:

1
2
search kube-system.svc.cluster.local svc.cluster.local cluster.local
options ndots:2

When we ping abc.123 (this domain has only one dot), the dns server’s log is below. Note the log tries abc.123.kube-system.svc.cluster.local. first, and only tries our domain last.

1
2
3
4
[INFO] 10.202.37.232:33386 - 36445 "A IN abc.123.kube-system.svc.cluster.local. udp 55 false 512" NXDOMAIN qr,aa,rd 148 0.001700129s
[INFO] 10.202.37.232:51389 - 58489 "A IN abc.123.svc.cluster.local. udp 43 false 512" NXDOMAIN qr,aa,rd 136 0.001117693s
[INFO] 10.202.37.232:32785 - 4976 "A IN abc.123.cluster.local. udp 39 false 512" NXDOMAIN qr,aa,rd 132 0.001047215s
[INFO] 10.202.37.232:57827 - 56555 "A IN abc.123. udp 25 false 512" NXDOMAIN qr,rd,ra 100 0.001763186s

Then when we ping abc.123.def (this domain has two dots), the dns server’s log looks like below. Note the log tries abc.123.def. first:

1
2
3
4
[INFO] 10.202.37.232:39314 - 794 "A IN abc.123.def. udp 29 false 512" NXDOMAIN qr,rd,ra 104 0.025049846s
[INFO] 10.202.37.232:51736 - 61456 "A IN abc.123.def.kube-system.svc.cluster.local. udp 59 false 512" NXDOMAIN qr,aa,rd 152 0.001213934s
[INFO] 10.202.37.232:53145 - 26709 "A IN abc.123.def.svc.cluster.local. udp 47 false 512" NXDOMAIN qr,aa,rd 140 0.001418143s
[INFO] 10.202.37.232:54444 - 1145 "A IN abc.123.def.cluster.local. udp 43 false 512" NXDOMAIN qr,aa,rd 136 0.001009799s

I hope this example makes two points clear:

  1. No matter what ndots is, the suffixes in the search parameter will be tried in order (our test used a nonexistent domain, so the resolver tried all possibilities)
  2. An improper ndots setting may pressure the dns server (if the domain exists, the dns query returns as soon as possible without continuing to search, reducing server pressure)

Optimization Discussion

Suppose ndots is now 2 and we want to query baidu.com. Since the dot count 1 is less than the configured 2, suffixes are appended for lookup first:

1
2
3
4
[INFO] 10.202.37.232:42911 - 55931 "A IN baidu.com.kube-system.svc.cluster.local. udp 57 false 512" NXDOMAIN qr,aa,rd 150 0.000116042s
[INFO] 10.202.37.232:53722 - 33218 "A IN baidu.com.svc.cluster.local. udp 45 false 512" NXDOMAIN qr,aa,rd 138 0.000075077s
[INFO] 10.202.37.232:46487 - 50053 "A IN baidu.com.cluster.local. udp 41 false 512" NXDOMAIN qr,aa,rd 134 0.000067313s
[INFO] 10.202.37.232:48360 - 51853 "A IN baidu.com. udp 27 false 512" NOERROR qr,aa,rd,ra 77 0.000127309s

Then we produce 3 useless dns query records. For the DNS server, just the domain baidu.com turns into 4x the traffic. What if n keeps growing — like the default 5 given in Kubernetes? We’d produce even more invalid requests, because not only baidu.com but even map.baidu.com and m.map.baidu.com would start trying from the search domains — quite a heavy pressure on the DNS server.

My personal suggestions:

  1. If requests between internal services are very frequent — i.e. we often access domains like xxx.svc.cluster.local — then keeping ndots larger is fine
  2. But when there are few requests between internal services, I strongly suggest lowering ndots to reduce useless traffic and lighten the dns server’s pressure For my personal use, changing it to 2 is good

Summary

Sorry that most of this article talks about how nameserver is resolved, with little on resolv.conf contents. The main reason: I’d been reading iptables the previous days, and this happened to come up, so I spent time on it — maybe with a bit of showing off.

When solving problems, understanding the parameters behind them is fairly important. I also pasted some of my experiments, hoping they help everyone — at least understand ndots before considering tuning.