The last blog introduced CronJob usage in Kubernetes and included suspend — but using this parameter well isn't so easy. This blog introduces the problems suspend brings, and gives solutions together with the source code.
The Problem Encountered
In the last blog, one problem brought by the suspend parameter was already mentioned:
If you change suspend from true to false — i.e. reopen the cron job —
missed tasks execute immediately (if no starting deadline is set);
K8S immediately schedules the previously missed tasks.
You may also encounter this error:
too many missed start time (> 100). Set or decrease .spec.startingDeadlineSeconds or check clock skew
The Meaning of startingdeadlineseconds – Official Documentation
The .spec.startingDeadlineSeconds field is optional. It stands for the deadline in seconds for starting the job if it misses its scheduled time for any reason. After the deadline, the cron job does not start the job. Jobs that do not meet their deadline in this way count as failed jobs. If this field is not specified, the jobs have no deadline.
The CronJob controller counts how many missed schedules happen for a cron job. If there are more than 100 missed schedules, the cron job is no longer scheduled. When .spec.startingDeadlineSeconds is not set, the CronJob controller counts missed schedules from status.lastScheduleTime until now.
For example, one cron job is supposed to run every minute, the status.lastScheduleTime of the cronjob is 5:00am, but now it's 7:00am. That means 120 schedules were missed, so the cron job is no longer scheduled.
If the .spec.startingDeadlineSeconds field is set (not null), the CronJob controller counts how many missed jobs occurred from the value of .spec.startingDeadlineSeconds until now.
For example, if it is set to 200, it counts how many missed schedules occurred in the last 200 seconds. In that case, if there were more than 100 missed schedules in the last 200 seconds, the cron job is no longer scheduled.
One thing to note from my own reading: even with suspend set back to false, a cron job that has missed more than 100 schedules will not run again.
A CronJob is counted as missed if it has failed to be created at its scheduled time. For example, If concurrencyPolicy is set to Forbid and a CronJob was attempted to be scheduled when there was a previous schedule still running, then it would count as missed.
If after reading the above you already understand suspend and startingDeadlineSeconds usage,
that’s excellent. If you still feel lost in the fog, follow me to read this part of the K8S source code.
Reading the cronjob Code in K8S:
I’ll use Kubernetes v1.17’s code as the example; the code below is excerpted from it, with all error handling removed:
// Run starts the main goroutine responsible for watching and syncing jobs. func(jm *Controller) Run(stopCh <-chanstruct{}) { // ... // the Run function syncs cron jobs every 10s go wait.Until(jm.syncAll, 10*time.Second, stopCh) // ... }
// syncAll lists all the CronJobs and Jobs and reconciles them. func(jm *Controller) syncAll() { // ... jobsBySj := groupJobsByParent(js) err = pager.New(pager.SimplePageFunc(cronJobListFunc)).EachListItem(context.Background(), metav1.ListOptions{}, func(object runtime.Object)error { sj, ok := object.(*batchv1beta1.CronJob) // sync each task syncOne(sj, jobsBySj[sj.UID], time.Now(), jm.jobControl, jm.sjControl, jm.recorder) cleanupFinishedJobs(sj, jobsBySj[sj.UID], jm.jobControl, jm.sjControl, jm.recorder) returnnil }) }
// get the time points of tasks not yet scheduled up to now; see the function below times, err := getRecentUnmetScheduleTimes(*sj, now)
// ... // the most recent time needing scheduling — you can see K8S only schedules the most recent task scheduledTime := times[len(times)-1] tooLate := false if sj.Spec.StartingDeadlineSeconds != nil { tooLate = scheduledTime.Add(time.Second * time.Duration(*sj.Spec.StartingDeadlineSeconds)).Before(now) } if tooLate { // if even the most recent scheduled time exceeds the deadline-allowed time, all tasks in the list are past deadline — stop executing, // and lastscheduletime is not set. klog.V(4).Infof("Missed starting window for %s", nameForLog) recorder.Eventf(sj, v1.EventTypeWarning, "MissSchedule", "Missed scheduled time to start a job: %s", scheduledTime.Format(time.RFC1123Z)) // TODO: Since we don't set LastScheduleTime when not scheduling, we are going to keep noticing // the miss every cycle. In order to avoid sending multiple events, and to avoid processing // the sj again and again, we could set a Status.LastMissedTime when we notice a miss. // Then, when we call getRecentUnmetScheduleTimes, we can take max(creationTimestamp, // Status.LastScheduleTime, Status.LastMissedTime), and then so we won't generate // and event the next time we process it, and also so the user looking at the status // can see easily that there was a missed execution. return } // concurrent task running forbidden — exit if sj.Spec.ConcurrencyPolicy == batchv1beta1.ForbidConcurrent && len(sj.Status.Active) > 0 { // Regardless which source of information we use for the set of active jobs, // there is some risk that we won't see an active job when there is one. // (because we haven't seen the status update to the SJ or the created pod). // So it is theoretically possible to have concurrency with Forbid. // As long the as the invocations are "far enough apart in time", this usually won't happen. // // TODO: for Forbid, we could use the same name for every execution, as a lock. // With replace, we could use a name that is deterministic per execution time. // But that would mean that you could not inspect prior successes or failures of Forbid jobs. klog.V(4).Infof("Not starting job for %s because of prior execution still running and concurrency policy is Forbid", nameForLog) return } // task replacement allowed — directly delete the previous active task if sj.Spec.ConcurrencyPolicy == batchv1beta1.ReplaceConcurrent { for _, j := range sj.Status.Active { klog.V(4).Infof("Deleting job %s of %s that was still running at next scheduled start time", j.Name, nameForLog)
// run this task jobReq, err := getJobFromTemplate(sj, scheduledTime) jobResp, err := jc.CreateJob(sj.Namespace, jobReq)
// Add the just-started job to the status list. ref, err := getRef(jobResp) sj.Status.Active = append(sj.Status.Active, *ref)
// update the schedule time sj.Status.LastScheduleTime = &metav1.Time{Time: scheduledTime}
return }
// getRecentUnmetScheduleTimes gets a slice of times (from oldest to latest) that have passed when a Job should have started but did not. // // If there are too many (>100) unstarted times, just give up and return an empty slice. // If there were missed times prior to the last known start time, then those are not returned. funcgetRecentUnmetScheduleTimes(sj batchv1beta1.CronJob, now time.Time) ([]time.Time, error) { // ... var earliestTime time.Time if sj.Status.LastScheduleTime != nil { earliestTime = sj.Status.LastScheduleTime.Time } else { // If none found, then this is either a recently created scheduledJob, // or the active/completed info was somehow lost (contract for status // in kubernetes says it may need to be recreated), or that we have // started a job, but have not noticed it yet (distributed systems can // have arbitrary delays). In any case, use the creation time of the // CronJob as last known start time. earliestTime = sj.ObjectMeta.CreationTimestamp.Time } if sj.Spec.StartingDeadlineSeconds != nil { // Controller is not going to schedule anything below this point schedulingDeadline := now.Add(-time.Second * time.Duration(*sj.Spec.StartingDeadlineSeconds))
if schedulingDeadline.After(earliestTime) { earliestTime = schedulingDeadline } } // ....
// earliestTime = MAX(creation time, last schedule time, current time - deadline seconds) // the loop below checks the time points that should have been scheduled during earliestTime => now, // and returns these time points; if the number of task points exceeds 100, error immediately for t := sched.Next(earliestTime); !t.After(now); t = sched.Next(t) { starts = append(starts, t) // An object might miss several starts. For example, if // controller gets wedged on friday at 5:01pm when everyone has // gone home, and someone comes in on tuesday AM and discovers // the problem and restarts the controller, then all the hourly // jobs, more than 80 of them for one hourly scheduledJob, should // all start running with no further intervention (if the scheduledJob // allows concurrency and late starts). // // However, if there is a bug somewhere, or incorrect clock // on controller's server or apiservers (for setting creationTimestamp) // then there could be so many missed start times (it could be off // by decades or more), that it would eat up all the CPU and memory // of this controller. In that case, we want to not try to list // all the missed start times. // // I've somewhat arbitrarily picked 100, as more than 80, // but less than "lots". iflen(starts) > 100 { // We can't get the most recent times so just return an empty slice return []time.Time{}, fmt.Errorf("too many missed start time (> 100). Set or decrease .spec.startingDeadlineSeconds or check clock skew") } } return starts, nil }
Summary
After reading Kubernetes’ cron job implementation, I basically have a general grasp of it.
The important point: K8S checks cron job time points every 10s,
checking unscheduled task time points within current time - DeadlineSeconds <=> current time,
selecting the last task to execute, and updating that task’s schedule time.
FAQ
The discussions below all assume the cron job’s minimum granularity is minutes!!!
In Linux cron tasks, * * * * * means executing once per minute.
Will Tasks Not Execute At All?
Possibly, in the following two situations:
Not setting DeadlineSeconds, or DeadlineSeconds too large, causing missed tasks to exceed 100 — K8S will not execute this cron job
When your DeadlineSeconds is set within 10s, each now - DeadlineSeconds may not include the task’s execution point
Will Tasks Execute More or Fewer Times?
Possibly, though at most one extra execution. K8S will execute it at the task’s next scheduling. Here’s the scenario:
Take the minimum granularity of one minute as an example — 1 minute >> 10s. Suppose one task execution took 2 minutes 22 seconds:
1 2 3 4 5 6 7 8
start end + + | | v v +----+----++-+-+ + + + ^ + + check
We’d trigger the check at 2 minutes 30 seconds; if DeadlineSeconds>30s,
the 2-minute point would be included in K8S’s scheduled tasks — a task that should have been ignored executes once.
This example also shows K8S cron jobs won’t execute tasks fewer times.
How Large Should DeadlineSeconds Be?
For now, setting Deadline below 6000s and above 10s works. I lean toward taking this time small —
within 10 minutes (600).
Anomalies to Handle
In the example above, the task’s actual execution time exceeded the task interval. A cron job that should execute every minute lost some tasks in the interval because one execution ran long.
len(times) > 1 indicates task backlog — we need to confirm the system status.