Millet Porridge

English version of https://corvo.myseu.cn

0%

CronJob Improvements in Kubernetes and Our Customization Needs

Earlier blogs introduced our use of K8s scheduled tasks and the source implementation of K8s cron jobs. But after actual use we found problems encountered during usage; I discuss solutions for each, hoping to help everyone, with suggestions appended at the end.

Some Usage of Cron Jobs in Kubernetes Reading the Kubernetes CronJob Source Code

Several Problems Encountered

  1. The existence of large numbers of scheduled tasks on machines makes docker’s burden heavy; in severe cases it even affects kernel speed. For the concrete phenomenon see Investigating a Kubernetes machine kernel problem

I don’t think this indicates a problem with K8s’s design; the design simply didn’t consider that docker’s performance on machines might be insufficient — unable to create containers in bulk quickly, slowing the whole system. We solved this problem via physical isolation: restricting scheduled tasks to a few fixed machines effectively reduces the probability of kernel problems on other cluster machines.

  1. Scheduled task run times are very inaccurate; some tasks’ execution gets delayed by minutes.

The delay problem doesn’t have a single cause; there are several types:

  1. K8s’s own scheduling delay — tasks that should start on time get dragged out for a long while
  2. Same as the reason above: docker’s burden on the machine is too heavy; containers that could start in seconds take half a minute longer. I’m not sure whether this problem appears in readers’ clusters, but in ours it’s especially obvious — Pods stay in ContainerCreating for a long time

20220227141906

K8s’s Improvements to Cron Jobs

In 2021 the CronJob API reached GA; an important change was replacing the cron job controller with v2. The original text is here:

https://kubernetes.io/blog/2021/04/09/kubernetes-release-1.21-cronjob-ga/

The original controller checked every 10 seconds whether all scheduled tasks needed execution; this operation could only be done by a single worker, with O(n) linear complexity — when there are too many scheduled tasks, performance becomes terrible. K8s introduced the new cron job controller in 1.19, changing the implementation strategy.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
// pkg/controller/cronjob/cronjob_controllerv2.go

// NewControllerV2 creates and initializes a new Controller.
func NewControllerV2(jobInformer batchv1informers.JobInformer, cronJobsInformer batchv1informers.CronJobInformer, kubeClient clientset.Interface) (*ControllerV2, error) {
jm := &ControllerV2{
// this queue is a delaying queue — items can be enqueued after a given delay
// t := nextScheduledTimeDuration(sched, now)
// jm.enqueueControllerAfter(curr, *t)
queue: workqueue.NewNamedRateLimitingQueue(workqueue.DefaultControllerRateLimiter(), "cronjob"),
recorder: eventBroadcaster.NewRecorder(scheme.Scheme, corev1.EventSource{Component: "cronjob-controller"}),

jobControl: realJobControl{KubeClient: kubeClient},
cronJobControl: &realCJControl{KubeClient: kubeClient},

jobLister: jobInformer.Lister(),
cronJobLister: cronJobsInformer.Lister(),

jobListerSynced: jobInformer.Informer().HasSynced,
cronJobListerSynced: cronJobsInformer.Informer().HasSynced,
now: time.Now,
}

// add hooks: cron job changes trigger notifications, allowing the controller to process tasks
cronJobsInformer.Informer().AddEventHandler(cache.ResourceEventHandlerFuncs{
AddFunc: func(obj interface{}) {
jm.enqueueController(obj)
},
UpdateFunc: jm.updateCronJob,
DeleteFunc: func(obj interface{}) {
jm.enqueueController(obj)
},
})
return jm, nil
}
// Run starts the main goroutine responsible for watching and syncing jobs.
func (jm *ControllerV2) Run(ctx context.Context, workers int) {
// multiple workers can be started for parallel processing
for i := 0; i < workers; i++ {
go wait.UntilWithContext(ctx, jm.worker, time.Second)
}
}

func (jm *ControllerV2) worker(ctx context.Context) {
for jm.processNextWorkItem(ctx) {
}
}

func (jm *ControllerV2) processNextWorkItem(ctx context.Context) bool {
// thanks to the delayed-enqueue mechanism, data taken from the queue is definitely a cron job needing execution now
key, quit := jm.queue.Get()
if quit {
return false
}
defer jm.queue.Done(key)

requeueAfter, err := jm.sync(ctx, key.(string))
switch {
case err != nil:
utilruntime.HandleError(fmt.Errorf("error syncing CronJobController %v, requeuing: %v", key.(string), err))
jm.queue.AddRateLimited(key)
case requeueAfter != nil:
jm.queue.Forget(key)
// postpone enqueue time to the next task execution
jm.queue.AddAfter(key, *requeueAfter)
}
return true
}

CronJob v2’s implementation leverages the K8s api server’s subscription-notification style:

  1. Classify the state types of cron jobs in etcd data into changing cron jobs and stably running cron jobs. Through this classification, executing cron jobs doesn’t require polling the whole list — only taking the tasks needing execution from the queue.
  2. The cron job execution function is distributed to multiple coroutines via the queue, effectively handling high-concurrency cron job problems.

The performance optimization after the update looks obvious

20220227143837

Our Improvements to Cron Jobs

Background

The above K8s cron job optimizations can’t be used by our cluster, because the cluster is fairly old and lacks this support. Another point: the above scheme only reduces task scheduling time — docker’s heavy burden problem remains unsolved. Given the heavy machine burden and inaccurate cron job execution times, we proposed a solution: lengthening the Job lifecycle of high-frequency cron jobs.

Scheme Design

For example, the user expects /bin/my_script to run once per minute. With our scheme, after starting the Pod we artificially keep the Pod alive for 1 hour or longer, adding cronjob scheduling inside the Pod to execute /bin/my_script once per minute. Of course the Pod’s lifetime is adjustable; we artificially set it to one hour so tasks spread across the running machines.

The original CronJob is as follows

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
apiVersion: batch/v1beta1
kind: CronJob
metadata:
name: hello
spec:
schedule: "* * * * *"
jobTemplate:
spec:
template:
spec:
containers:
- name: hello
image: busybox
args:
- /bin/sh
- -c
- date; echo Hello from the Kubernetes cluster
restartPolicy: OnFailure

The modified CronJob is as follows

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
apiVersion: batch/v1beta1
kind: CronJob
metadata:
name: hello
spec:
# lower the run frequency
schedule: "0 * * * *"
jobTemplate:
spec:
template:
spec:
containers:
- name: hello
image: busybox
args:
- /bin/do-cron # via our own script, create the cronjob and start crond
- /bin/sh
- -c
- date; echo Hello from the Kubernetes cluster
env:
- name: CRON_SCHEDULE # pass the original cron into the container via an environment variable
value: "* * * * *"
restartPolicy: OnFailure

Problems with the Scheme and How to Solve Them

The scheme’s benefits:

  1. Machine burden greatly reduced: creating 60 Pods per hour becomes 1 Pod per hour
  2. Cron job run timing is more accurate: single-machine tasks running every minute have basically no error — very friendly for cron jobs needing fine control

Some problems this brings:

  1. Lengthening the Pod lifecycle: each Pod startup, the previous Pod may already be closed or not yet closed, causing task loss or duplication
  2. Lengthening the Pod lifecycle: each Pod may execute multiple tasks in parallel, making resource control less precise

For the first problem, we can avoid it through certain mechanisms; for the second, due to the design itself, there’s no good solution. In actual use, the high-frequency cron jobs we encountered aren’t very resource-sensitive.

How to Ensure Cron Job Availability and Stability

This part involves implementation details; I’ll only introduce some logic without concrete code. Two aspects need consideration:

  1. How to seamlessly connect cron job execution, ensuring no loss or duplication
  2. After users modify tasks or deploy new versions, how to refresh and update cron jobs as quickly as possible

Container Redundancy

20220227171038

  • Not losing tasks:

After the new Pod starts, the old Pod doesn’t immediately go offline; we provide it a small buffer interval — as shown, the dashed region on the timeline where both Pods run simultaneously. Designed this way, we can guarantee no task loss.

  • Not duplicating tasks:

Each of our containers has the concept of a container token. When Pod1 runs, it holds the token; after we start Pod2, Pod1 releases the token at a suitable moment, and Pod2 can execute cron jobs only after obtaining the token. The timing of releasing and obtaining tokens also matters: for Pod1 we start releasing at the 10th second after a certain minute begins — i.e. releasing the token in the first half of the minute — so Pod2 has about 50s to obtain the token. This time is ample, enough for Pod2 to get the application token and start executing the next task.

Separated Execution

During cron job execution, users may well switch versions or modify cron jobs at non-hour-boundary times. Once that happens, the container redundancy above guarantees we update in the next scheduling cycle — but when users modify tasks or launch versions, they want it to take effect immediately, not wait (possibly an hour). Based on this idea, we considered a way to separate ordinary cron jobs from manually changed tasks; here’s the concrete logic diagram:

20220227171627

The implementation here mainly uses one K8s cron job feature: kubectl create job --from=cronjob/<cronjob-name> <job-name> Manually created scripts also acquire the token; Pod1 ends early, and until Pod2 starts running, the Manual Pod bears the task of running scripts. The idea here is separating routine behavior from ad-hoc behavior.

Suggestions for Using Cron Jobs

  1. Determine the cron job scale: once per hour or once per minute
  2. Determine the tolerance for cron job run delay: can you accept cron jobs running minutes late
  3. Physically isolate cron job machines: even using our own strategy, with each cron job Pod’s lifecycle lengthened, we found cron machines’ io usage very high — I suggest adding SSDs directly to such machines.
  4. Pay attention to logging and related alarms

Summary

I’ve only roughly introduced our cron job optimization; there are many concrete details — especially the cron job monitoring code, which is even more than its implementation code. Our strategy has run online for over a year and should be a fairly stable feature now, so I share the design strategy for everyone’s reference.

Sometimes we use certain frameworks just because they’re handy, but as the business develops, frameworks need gradual customization and optimization to fit business needs. Sustainably solving business development needs is what effectively advances K8s component adoption.

Once you use open source components, be aware: you need to customize certain strategies yourself to solve problems, and no one will take responsibility for any of your actions. You can read this thread: My self-built Gitlab opened to the public network got hacked.