Millet Porridge

English version of https://corvo.myseu.cn

0%

Adding a Program Profiling Feature to a Kubernetes-based PaaS Platform

In early March we considered adding performance profiling and analysis tools to the Kubernetes cluster (mainly targeting Python, especially uwsgi applications). After researching several profiling schemes, we chose py-spy for uwsgi application profiling. This blog introduces the concrete implementation flow and some debugging strategies; the appendix introduces what I learned.

Profile Tools

Our cluster mainly profiles and analyzes Python applications. Our original non-kubernetes scheme was pyflame, but it’s no longer maintained. Another colleague suggested py-spy — still under fairly active development, with more features than pyflame and simpler installation — so we switched to py-spy.

System Implementation Scheme

Running Directly

I quite like its top feature — you can immediately confirm current stack information.

Its README also introduces how to run py-spy in docker containers: PTRACE permission must be added, and hostPID must also be readable — if you can only see your own pid inside the container, sampling is impossible.

Running in a kubernetes Environment

The Scheme Referenced from github

Below is pyflame’s implementation scheme: https://github.com/monsterxx03/kube-pyflame/blob/master/kubectl-pyflame

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
nodeName=$(kubectl get pod ${POD} -n ${NAMESPACE} -o jsonpath='{.spec.nodeName}')
kubectl create -f - <<EOF
apiVersion: v1
kind: Pod
metadata:
name: ${DEBUG_POD}
namespace: ${NAMESPACE}
spec:
terminationGracePeriodSeconds: 0
hostPID: true
containers:
- command: ['sh', '-c', 'while true; do sleep 10; done;']
image: ${IMAGE}
imagePullPolicy: IfNotPresent
name: pyflame
securityContext:
privileged: true
nodeSelector:
kubernetes.io/hostname: ${nodeName}
EOF
kubectl wait --for=condition=Ready pod/${DEBUG_POD} -n ${NAMESPACE} --timeout=120s
kubectl exec ${DEBUG_POD} -n ${NAMESPACE} -- bash -c "pyflame -p ${pid} -s $SECONDS -r $RATE | flamegraph.pl > /tmp/pyflame.svg"

This is the simplest implementation. Two points to note: the hostPID: true and securityContext: privileged: true here — these grant special permissions.

Our Implementation

Based on the pyflame method above, using py-spy is about the same — the configuration is very similar, but some places differ:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
---
apiVersion: batch/v1
kind: Job
metadata:
name: profile-task-22-1618198229
namespace: default
spec:
ttlSecondsAfterFinished: 259200
backoffLimit: 3
template:
metadata:
labels:
app_id: "11"
spec:
nodeSelector:
kubernetes.io/hostname: "minikube"
restartPolicy: Never
hostPID: true
containers:
- name: profile-task--22
image: <IMAGE>
args:
- profile_task
env:
- name: params
value: "{\"rate\": 100, \"seconds\": 60, \"exclude_idle\": true, \"multi_threads\": true, \"app_name\": \"arya-python3\", \"version_id\": 2255}"
securityContext:
capabilities:
add:
- SYS_PTRACE

The differences are as follows:

  1. We use a job instead of a pod — mainly to standardize the code; also, jobs have this parameter: ttlSecondsAfterFinished
  2. I default the namespace to default rather than the application’s namespace. The main reason: in an earlier article I introduced: A Solution for Providing Dashboard Support on a Kubernetes-based PaaS Platform. There we opened operation permissions on users’ own application namespaces to users — i.e. they can exec into containers. That creates a problem: because this sampling container has hostPID permission, if ordinary users could enter this privileged container, they could kill other programs — an operation that cannot be allowed. Therefore such privileged containers must be uniformly placed in namespaces users cannot access.
  3. nodeSelector specifies the machine on which the sampling task must run; this machine is randomly selected by the PaaS — users needn’t care which machine their application is on
  4. The environment variables store the sampling parameters; after the task starts running, sampling begins automatically, and after sampling ends the result is uploaded — users get the data with a few button clicks

Code Details in the Sampling Container

This program does several jobs:

  1. Read the configuration from environment variables
  2. Confirm the process id to sample
  3. Sample
  4. Report the flame graph

I’ll paste the main code; error handling is omitted:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
def _run_profile_task():
params = json.loads(os.environ['params'])
app_name = params['app_name']
version_id = params['version_id']

# query the corresponding process info — here we search the uwsgi worker name,
# because we set uwsgi startup parameters adding a process name prefix: procname-prefix: arya-python3-2255_
proc_name = '%s-%s_uWSGI worker' % (app_name, version_id)
for proc in psutil.process_iter():
cmdline = proc.cmdline()
if len(cmdline) > 0:
cmd = " ".join(cmdline)
if cmd.startswith(proc_name):
target_proc = proc
break
output_file = os.path.join(
'var', 'run',
'profile-{}-{}-{}'.format(app_name, version_id, int(time.time()))
)
proc_args = [
'py-spy', 'record', '--subprocesses',
'--duration', '{}'.format(params['seconds']),
'--rate', '{}'.format(params['rate']),
'--pid', '{}'.format(target_proc.pid),
'--output', output_file,
]
if params.get('multi_threads'):
pass
if params.get('exclude_idle', True) is False:
proc_args.append('--idle')

# run the sampling process
_ret, _out, _err = run_with_timeout(proc_args, get_timeout())

# report...

Program Security Considerations

The security issue is mainly the namespace restriction strategy — the concrete reasons and scheme were given in the implementation comparison above.

Summary

This article mainly wanted to introduce the profiling tool we added in a kubernetes environment according to our PaaS platform’s situation. Currently it stops at Python applications, but extending to other languages’ programs in future is easy. Another point to note is handling privileged containers: grant as few permissions as possible, and don’t let users see them.

Appendix 1: Job Debugging Strategy

In pyflame’s implementation scheme there’s a line of configuration worth learning — simply a debugging godsend for job tasks and pod containers:

1
- command: ['sh', '-c', 'while true; do sleep 10; done;']

Appendix 2: Learning the kubectl flame Tool

After our feature went live, I also saw this sampling tool — provided via the kubectl plugin development approach. If you only need a simple sampling tool, there’s no need to develop a whole system; try this kubectl plugin:

Introducing Kubectl Flame: Effortless Profiling on Kubernetes

It contains this sentence: Profiling is a non-trivial task.

1
kubectl flame mypod -t 1m -f /tmp/flamegraph.svg

I also read through the code: https://github.com/VerizonMedia/kubectl-flame The codebase isn’t long; let me briefly introduce its implementation.

Code Implementation

cli/cmd/kubernetes/root.go — the entry code is here:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
cmd := &cobra.Command{
Use: "flame [pod-name]",
DisableFlagsInUseLine: true,
Short: "Profile running applications by generating flame graphs.",
Long: flameLong,
Example: fmt.Sprintf(flameExamples, "kubectl"),
PersistentPreRun: func(c *cobra.Command, args []string) {
c.SetOutput(streams.ErrOut)
},
Run: func(cmd *cobra.Command, args []string) {
// check whether the language and event are supported
validateFlags(chosenLang, chosenEvent, &targetDetails, &jobDetails);

targetDetails.PodName = args[0]
if len(args) > 1 {
targetDetails.ContainerName = args[1]
}

cfg := &data.FlameConfig{
TargetConfig: &targetDetails,
JobConfig: &jobDetails,
ConfigFlags: options.configFlags,
}

Flame(cfg)
},
}

The Flame function is the concrete sampling function; I’ve omitted all error handling:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
func Flame(cfg *data.FlameConfig) {
ns, err := kubernetes.Connect(cfg.ConfigFlags)
p := NewPrinter(cfg.TargetConfig.DryRun)

cfg.TargetConfig.Namespace = ns
ctx := context.Background()

// handle the info of the pod to listen to
p.Print("Verifying target pod ... ")
pod, err := kubernetes.GetPodDetails(cfg.TargetConfig.PodName, cfg.TargetConfig.Namespace, ctx)
containerName, err := validatePod(pod, cfg.TargetConfig)

containerId, err := kubernetes.GetContainerId(containerName, pod)
p.PrintSuccess()

cfg.TargetConfig.ContainerName = containerName
cfg.TargetConfig.ContainerId = containerId

p.Print("Launching profiler ... ")
// launch the profile task
profileId, job, err := kubernetes.LaunchFlameJob(pod, cfg, ctx)
if cfg.TargetConfig.DryRun {
return
}

cfg.TargetConfig.Id = profileId
profilerPod, err := kubernetes.WaitForPodStart(cfg.TargetConfig, ctx)
p.PrintSuccess()
apiHandler := &handler.ApiEventsHandler{
Job: job,
Target: cfg.TargetConfig,
}
done, err := kubernetes.GetLogsFromPod(profilerPod, apiHandler, ctx)
<-done
}

I won’t paste the later code — it’s just for launching a Job. Fast-forward directly to the concrete profile code:

1
2
3
4
5
6
7
// agent/profiler/python.go
cmd := exec.Command(pySpyLocation, "record", "-p", pid, "-o", pythonOutputFileName, "-d", duration, "-s", "-t")
var out bytes.Buffer
var stderr bytes.Buffer
cmd.Stdout = &out
cmd.Stderr = &stderr
// honestly, these parameters are fairly simple

How Applications in Different Languages Should Be Profiled

Also learned from reading the codebase…

Language Profiling tool
Java async-profiler
Python py-spy
Golang bcc-profiler
Ruby rbspy