Millet Porridge

English version of https://corvo.myseu.cn

0%

A Bizarre K3s DNS Error Investigation: When K3s Meets fnNAS

In daily Kubernetes troubleshooting, network problems often consume the most time. This time I ran a K3s (single-node master) via Docker on an fnNAS, preparing to run argocd-diff-preview for CI. After Argo CD finished deploying, it kept reporting DNS resolution failures; the logs looked like “querying DNS with localhost inside the Pod” — very counterintuitive.

The root cause finally located was also very “invisible”: the host directory’s Default ACL affected the resolv.conf that containerd generates and mounts for Pods, so non-root containers couldn’t read /etc/resolv.conf, triggering the libc resolver’s fallback logic — blindly querying 127.0.0.1/[::1]:53.

Background

I use Argo CD locally to manage and release applications. This time I wanted to introduce argocd-diff-preview for CI checks, writing diff results back to PRs as GitHub comments.

The effect is roughly like this:

1774962665949.png

To run such tasks I needed a lightweight Kubernetes cluster. I happened to have an fnNAS at hand, so I started a K3s server on it with Docker (servicelb disabled); the docker-compose is:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
version: "3.8"
services:
argocd-k3s:
image: k3s:v1.35.1-k3s1
container_name: argocd-k3s
privileged: true
network_mode: bridge # prevent interference from the host Docker network's DNS
restart: unless-stopped
ports:
- "9444:6443"
command:
- "server"
- "--prefer-bundled-bin"
- "--disable=servicelb"
- "--node-name=argocd-k3s"
- "--kube-apiserver-arg=event-ttl=10m"
volumes:
- /vol1/1000/K3sData/argocd-k3s/data:/var/lib/rancher/k3s
- /dev:/dev:ro

Then executing argocd-diff-preview in CI started erroring:

1
2
3
4
5
6
7
8
9
10
11
argocd-diff-preview  \
--create-cluster=False --keep-cluster-alive \
-d \
--argocd-chart-url ${chart_url} \
--repo ${repo} \
--target-branch ${env.CHANGE_BRANCH} \
--argocd-chart-name argocd-apps

Tue, 31 Mar 2026 12:42:20 UTC DBG AppsetGenerateWithRetry attempt 1/5 to Argo CD...
Tue, 31 Mar 2026 12:42:21 UTC DBG Waiting 1s before next appset generate attempt (2/5)...
Tue, 31 Mar 2026 12:42:21 UTC WRN Appset generate attempt 1/5 failed. error="failed to run argocd appset generate: argocd command failed: {\"level\":\"fatal\",\"msg\":\"rpc error: code = Internal desc = unable to resolve git revision : failed to list refs: dial tcp: lookup argocd-redis on [::1]:53: read udp [::1]:59984-\\u003e[::1]:53: read: connection refused\",\"time\":\"2026-03-31T12:42:21Z\"}\n: exit status 20"

The Phenomenon: Querying DNS via [::1]:53 Inside Pods

Argo CD couldn’t sync applications normally. Checking argocd-server logs revealed many similar errors:

1
Failed to resync revoked tokens... dial tcp: lookup argocd-redis on [::1]:53: read udp [::1]:34952->[::1]:53: read: connection refused

My first reaction seeing this log: why would a Pod query DNS via the IPv6 loopback address ([::1])?

Following conventional thinking I first checked CoreDNS, but CoreDNS was running perfectly. That made it even stranger: a Pod with dnsPolicy: ClusterFirst should theoretically send queries to CoreDNS (e.g. 10.43.0.10), not retry against local port 53.

Investigation: Starting from /etc/resolv.conf

Since nothing looked wrong from outside, I entered the container directly to confirm the DNS configuration:

1
kubectl -n argocd exec -ti argocd-server-fbbdf5d4d-c8z5q -- /bin/bash

Then viewed /etc/resolv.conf:

1
2
argocd@argocd-server:~$ cat /etc/resolv.conf
cat: /etc/resolv.conf: Permission denied (os error 13)

At this point the logs could basically be explained: the application process couldn’t read /etc/resolv.conf, so it got no nameserver configuration.

In this situation the libc resolver takes its fallback logic, directly trying to send DNS queries to local localhost (127.0.0.1 and [::1]); if no DNS service listens on local port 53, you get connection refused.

So this wasn’t an IPv6 problem per se — the container simply “couldn’t see” the correct nameserver.

Root Cause: The “+” at the End of File Permissions (ACL)

Continuing with ls -l to view permissions:

1
2
argocd@argocd-server:~$ ls -alh /etc/resolv.conf 
-rw-r-----+ 1 root root 102 Mar 31 11:56 /etc/resolv.conf

After I pasted this content to the AI, it mentioned the ACL issue:

Note the + at the end of the permissions: it indicates the file has ACL (Access Control List) enabled. Even if base permissions look fine, the ACL may additionally tighten “other” users’ permissions. Since Argo CD runs as a non-root user by default (e.g. UID 999), it gets shut out.

/etc/resolv.conf isn’t a file in the image — it’s mounted in at runtime. To confirm which file it corresponds to on the host, I went back to the host and used crictl inspect to check the container mount information:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
~ # crictl ps  # find the argocd-server container's ID
CONTAINER IMAGE CREATED STATE NAME ATTEMPT POD ID POD NAMESPACE
c2722b3817277 86219ff1fefea 2 seconds ago Running argocd-server 3 d32169958ed68 argocd-server-fbbdf5d4d-tjsnf argocd

~ # crictl inspect <container_id> # view /etc/resolv.conf's real source
# {
# "destination": "/etc/resolv.conf",
# "options": [
# "rbind",
# "rprivate",
# "ro"
# ],
# "source": "/var/lib/rancher/k3s/agent/containerd/io.containerd.grpc.v1.cri/sandboxes/fd2373cfd801e2f95267286af155df3ca0cd912563bc9a668cc5e1b3a4956d3e/resolv.conf",
# "type": "bind"
# }

~ # cat /var/lib/rancher/k3s/agent/containerd/io.containerd.grpc.v1.cri/sandboxes/fd2373cfd801e2f95267286af155df3ca0cd912563bc9a668cc5e1b3a4956d3e/resolv.conf
search argocd.svc.cluster.local svc.cluster.local cluster.local
nameserver 10.43.0.10
options ndots:5

Since my K3s data directory is on /vol1/1000/K3sData (a NAS mounted disk), I followed this path over and used getfacl to view the ACL:

1
2
3
4
5
6
root@Host:~# getfacl /vol1/.../sandboxes/.../resolv.conf
# owner: root
# group: root
user::rw-
group::--x
other::--- <--- the real culprit is here!

The truth emerged: NAS mounted disks like /vol1 may carry strict Default ACLs by default. When containerd dynamically generates resolv.conf for Pods under this directory, it inherits the parent directory’s Default ACL, tightening other permissions to ---. Non-root processes in the Pod then can’t read /etc/resolv.conf, triggering the “query localhost DNS” fallback behavior.

Solution: Cleaning the Default ACL Inheritance Rules

I ultimately didn’t change K3s configuration, nor disable ACLs on the whole disk (which usually affects services like Samba). The more reliable approach: clean only the Default ACL rules causing inheritance on the K3s/containerd working directory.

Execute on the host:

1
2
3
4
5
# 1. clear the Default ACL inheritance rules causing the problem on the sandbox directory
sudo setfacl -k /vol1/1000/K3sData/

# 2. (optional) batch-reset base permissions for currently existing directories
sudo setfacl -R -b /vol1/1000/K3sData/

Afterwards delete Argo CD’s Pods to let them rebuild. After the new Pods start, the resolv.conf generated by containerd returns to normal readable permissions; entering the container and running cat /etc/resolv.conf no longer reports Permission denied. At this point Argo CD immediately returned to normal, and applications could Sync normally.

Automatic Persistence: Using tmpfiles.d (advanced scheme)

If you worry the NAS system will re-“inherit” the wrong ACL rules after reboot or Web UI operations, you can use systemd’s tmpfiles.d mechanism — a more elegant “declarative” configuration than Crontab scripts.

Create the config file /etc/tmpfiles.d/k3s-nas-acl.conf on the host:

1
2
3
4
5
6
# ensure the directory and its subdirectories' base permissions are 0755
# type path mode user group age argument
z /vol1/1000/K3sData/ 0755 root root - -

# recursively force-apply rules: ensure others can read, and newly generated children inherit this permission
A+ /vol1/1000/K3sData/ - - - - other::r-x,default:other::r-x

After configuring, run the following command to take effect immediately (no reboot needed):

1
sudo systemd-tmpfiles --create /etc/tmpfiles.d/k3s-nas-acl.conf

The system scans the path and automatically corrects all nonconforming ACL entries per the configuration.

Summary and Pitfall-Avoidance Guide

  • Don’t be misled by [::1]:53: when a Pod queries localhost DNS, first check whether it can read /etc/resolv.conf.
  • Mind non-root containers: more and more components run as non-root by default; files generated/mounted by the host must guarantee “other readable” or at least readable by the application user.
  • Beware NAS default ACLs: when running K3s/Docker on NAS mounted disks, the + at the end of ls -l permissions is often the signal of permission problems; for bizarre phenomena, locate with getfacl first.
  • AI is really too powerful — asking the right questions makes debugging more efficient.

I hope this record gives you a faster investigation entry point.