To monitor a Kubernetes CronJob, have the job ping a heartbeat URL when it finishes successfully, and alert when that ping does not arrive on schedule. kubectl and kube-state-metrics show what the CronJob controller did, but a run that is skipped, never scheduled, or stuck behind a pod that never starts produces no failed Job, so nothing in the cluster turns red.
I'm Léo, I build Hyperping. This guide covers the CronJob failure modes that stay silent, the native commands to inspect them, and the setup I recommend with Hyperping healthchecks. I tested every manifest and command below on Kubernetes 1.37.1 (the current release in October 2026) and the container script in an Alpine 3.24 root filesystem.
Key takeaways
- A CronJob with
concurrencyPolicy: Forbidwhose pod is stuck inPendingorImagePullBackOffblocks every following run. The only trace is aJobAlreadyActiveevent, andkubectl get jobsshows the Job asRunning. - A pod that never starts does not count toward
backoffLimit. OnlyactiveDeadlineSecondsturns it into a failed Job. - Use
spec.timeZonewith an IANA name.CRON_TZ=insidescheduleis rejected, and withouttimeZonethe schedule follows the controller's clock, usually UTC. - After a controller outage, a CronJob runs once at most for all the schedules it missed, and only if
startingDeadlineSecondsallows it. - An external heartbeat catches all of these with one rule: no success ping by the expected time plus the grace period means an alert.
How Kubernetes CronJobs fail silently
The CronJob docs are upfront about it: a CronJob creates a Job "approximately once per execution time of its schedule", and "two Jobs might be created, or no Job might be created". Most of the gap between "approximately" and "always" is invisible.
A pod that never starts blocks the schedule
With concurrencyPolicy: Forbid, "if it is time for a new Job run and the previous Job run hasn't finished yet, the CronJob skips the new Job run". That is the right setting for backups and migrations, but it turns one bad run into a stalled schedule.
Push an image tag that does not exist, or ask for more memory than any node has, and the pod sits in Pending. It never fails, so it never counts against backoffLimit (6 by default), and the Job never finishes. I reproduced it on a 1.37.1 control plane with no schedulable node:
$ kubectl get jobs
NAME STATUS COMPLETIONS DURATION AGE
stuck-29858136 Running 0/1 2m5s 2m5s
$ kubectl get events --field-selector involvedObject.name=stuck \
-o custom-columns=REASON:.reason,MESSAGE:.message
REASON MESSAGE
SuccessfulCreate Created job stuck-29858136
JobAlreadyActive Not starting job because prior execution is running and concurrency policy is ForbidThe Job says Running while its pod is Pending. Every following schedule produces another JobAlreadyActive event and no Job, until someone deletes the stuck Job or activeDeadlineSeconds fails it.
With the default concurrencyPolicy: Allow, the same problem piles up Jobs instead: each run creates a new Pending pod.
Missed schedules are dropped, not replayed
When the controller cannot start a run on time (control plane upgrade, controller restart, suspended CronJob), it does not replay every missed run. It creates one Job for the most recent missed time, and only if that time is still within startingDeadlineSeconds. Older runs are gone, and a run skipped for being past the deadline leaves a MissSchedule warning event.
The docs also describe a limit at 100 missed schedules: "it does not start the Job and logs the error too many missed start times. Set or decrease .spec.startingDeadlineSeconds or check clock skew". That was the behavior of the old controller, which stopped scheduling. The current controller, the default since 1.21, emits a TooManyMissedTimes warning event and still creates the most recent Job. Setting startingDeadlineSeconds keeps the count small in both cases, because the controller then only counts misses inside that window.
The schedule runs in the controller's timezone
Without spec.timeZone, "the kube-controller-manager interprets schedules relative to its local time zone". Control planes usually run in UTC, so 0 2 * * * runs at 2:00 UTC, not 2:00 in your office. Check with kubectl get cronjob: the TIMEZONE column shows <none> when the field is missing.
timeZone takes an IANA name and has been stable since 1.27. The old trick of prefixing the schedule fails validation:
The CronJob "tz-in-schedule" is invalid: spec.schedule: Invalid value:
"CRON_TZ=Europe/Paris 0 2 * * *": cannot use TZ or CRON_TZ in schedule, use timeZone field insteadtimeZone: Local is rejected the same way: "timeZone must be an explicit time zone".
History limits delete the evidence
By default a CronJob keeps 3 successful Jobs and 1 failed Job. If a job fails twice in a row overnight, the first failure and its pod logs are deleted before you look. Raise failedJobsHistoryLimit if you debug from pod logs.
How to inspect CronJobs with kubectl
These commands answer "did it run, and why not":
# Schedule, timezone, suspended, active Jobs and last schedule time
kubectl get cronjob nightly-backup
# Jobs as they are created and finish
kubectl get jobs --watch
# Why runs were skipped: JobAlreadyActive, MissSchedule, TooManyMissedTimes
kubectl get events --field-selector involvedObject.kind=CronJob,involvedObject.name=nightly-backup
# Last scheduled and last successful run, in UTC
kubectl get cronjob nightly-backup -o jsonpath='{.status.lastScheduleTime}{"\n"}{.status.lastSuccessfulTime}{"\n"}'
# Run it now from the same template
kubectl create job backup-test-1 --from=cronjob/nightly-backupkubectl get cronjob prints NAME SCHEDULE TIMEZONE SUSPEND ACTIVE LAST SCHEDULE AGE. An ACTIVE count that never drops back to 0 is the stuck pod from above.
If you run Prometheus, kube-state-metrics exposes kube_cronjob_status_last_successful_time and kube_cronjob_next_schedule_time, and kube_job_status_failed for Jobs. An alert on time() - kube_cronjob_status_last_successful_time works, but it runs in the same cluster as the job: a broken node pool, a Prometheus that ran out of memory, or an Alertmanager with a bad route silences it too. An external check does not share those failures. For the wider cluster picture, see my Kubernetes monitoring setup guide.
How to monitor Kubernetes CronJobs with Hyperping
A Hyperping healthcheck gives the job a secret URL. When no ping arrives by the scheduled time plus a grace period, Hyperping opens an incident and alerts you. It resolves on the next successful ping. Healthchecks are included on every plan, Free included.
1. Create a healthcheck with the same schedule and timezone
In Hyperping, open Healthchecks, click Create healthcheck, and pick Cron. Copy spec.schedule into Cron expression and pick the spec.timeZone value in Timezone: 0 2 * * * and Europe/Paris for the example below. If the CronJob has no timeZone, pick the controller's timezone, usually UTC, or better, add timeZone to the CronJob.
Use the cron expression generator to check the next run times, for example every 6 hours or every day at midnight.
2. Store the ping URL in a Secret
apiVersion: v1
kind: Secret
metadata:
name: hyperping
type: Opaque
stringData:
backup-url: https://hc.hyperping.io/tok_your_backup_tokenThe token is the only credential: anyone with the URL can send pings for that healthcheck, so treat it like a password.
3. Ping /start, then ping only on success
If the image has a shell and wget or curl, wrap the command. Alpine and BusyBox images have wget (BusyBox's) but no curl. Debian slim images have neither.
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-backup
spec:
schedule: "0 2 * * *"
timeZone: "Europe/Paris"
concurrencyPolicy: Forbid
startingDeadlineSeconds: 600
jobTemplate:
spec:
backoffLimit: 2
activeDeadlineSeconds: 3600
template:
spec:
restartPolicy: Never
containers:
- name: backup
image: ghcr.io/acme/backup:1.4.2
env:
- name: HC_URL
valueFrom:
secretKeyRef:
name: hyperping
key: backup-url
command: ["/bin/sh", "-c"]
args:
- |
wget -q -T 10 -O /dev/null "$HC_URL/start" || true
/app/backup.sh || exit $?
wget -q -T 10 -O /dev/null "$HC_URL" || echo "Hyperping ping failed" >&2Each line has a job:
/startopens a run in Hyperping so it can record the duration.|| truekeeps a network hiccup from cancelling the backup.|| exit $?stops on failure with the job's own exit code, so Kubernetes marks the pod failed and retries it withinbackoffLimit. No success ping is sent.- The success ping is allowed to fail without failing the pod. Otherwise an unreachable endpoint would make Kubernetes run a successful backup a second time.
I ran this script in an Alpine 3.24 root filesystem against a local listener: a backup exiting 0 sent /start then the success ping, a backup exiting 3 sent only /start and the shell exited 3, and with the listener down the shell still exited 0. Against hc.hyperping.io, BusyBox wget handled HTTPS without extra packages.
4. Use init containers when the image has no shell
Distroless and scratch images have no shell and no HTTP client. Init containers run in order and must each succeed before the regular containers start, and with restartPolicy: Never a failed init container fails the pod. So run the job as an init container and the success ping as the main container:
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-backup
spec:
schedule: "0 2 * * *"
timeZone: "Europe/Paris"
concurrencyPolicy: Forbid
startingDeadlineSeconds: 600
jobTemplate:
spec:
backoffLimit: 2
activeDeadlineSeconds: 3600
template:
spec:
restartPolicy: Never
initContainers:
- name: ping-start
image: curlimages/curl:8.22.0
env:
- name: HC_URL
valueFrom:
secretKeyRef:
name: hyperping
key: backup-url
command: ["sh", "-c", 'curl -fsS -m 10 --retry 3 -o /dev/null "$HC_URL/start" || true']
- name: backup
image: ghcr.io/acme/backup:1.4.2
containers:
- name: ping-success
image: curlimages/curl:8.22.0
env:
- name: HC_URL
valueFrom:
secretKeyRef:
name: hyperping
key: backup-url
command: ["sh", "-c", 'curl -fsS -m 10 --retry 3 -o /dev/null "$HC_URL" || echo "Hyperping ping failed" >&2']If backup exits non-zero, ping-success never starts and the healthcheck goes down once the grace period passes.
Both manifests pass kubectl apply --dry-run=server on 1.37.1 and kubeconform -strict.
5. Set the CronJob guardrails
The heartbeat tells you something is wrong. These fields keep a single bad run from blocking the next ones:
| Field | Value in the example | Effect |
|---|---|---|
concurrencyPolicy |
Forbid |
Never two backups at once. A skipped run sends no ping, so you hear about it. |
startingDeadlineSeconds |
600 |
A run delayed by more than 10 minutes is skipped (MissSchedule) instead of starting at a random time. |
activeDeadlineSeconds |
3600 |
A Job still active after an hour, including one stuck in Pending, fails with DeadlineExceeded and frees the schedule. |
backoffLimit |
2 |
Two retries for a job that crashes, instead of 6. |
I checked the activeDeadlineSeconds behavior with a pod that could never be scheduled: the Job was marked Failed with reason DeadlineExceeded, and the CronJob created the next run on schedule.
6. Size the grace period and route the alerts
In cron mode, Hyperping expects the success ping by the scheduled time plus the grace period. For a CronJob, that window has to cover the start delay (up to startingDeadlineSeconds), the image pull, and the run itself. A backup that takes 20 minutes with a 10 minute starting deadline needs at least 30 minutes of grace. I would set 45.
Once a few runs have gone through, the healthcheck's ping history shows each run's duration between /start and the success ping. Adjust from there.
When a run is missed, Hyperping alerts every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so route them to PagerDuty or Opsgenie when a failed backup has to wake the person on call.
Test the whole chain
- Run
kubectl create job backup-test-1 --from=cronjob/nightly-backupand check the pings in the healthcheck's last pings list (user agentWgetorcurl). - Break it on purpose: point the image at a tag that does not exist and run the same command. The pod stays in
ImagePullBackOff, no success ping arrives, and the alert reaches you after the grace period. - Restore the tag and run it again. The next success ping closes the incident.
Other schedulers fail in their own ways, and the same heartbeat covers them: the Laravel scheduler, GitHub Actions scheduled workflows, systemd timers and Windows Task Scheduler. To compare heartbeat tools, read the best cron job monitoring tools. If the cluster runs Prometheus, a dead man's switch for Alertmanager tells you when the alerting pipeline itself goes quiet.
FAQ
Why is my Kubernetes CronJob not running? ▼
Run `kubectl get cronjob` and `kubectl get events --field-selector involvedObject.kind=CronJob`. The usual causes are a suspended CronJob, a previous Job still active with `concurrencyPolicy: Forbid` (event `JobAlreadyActive`), a run skipped because it started after `startingDeadlineSeconds` (event `MissSchedule`), or a schedule read in the controller's timezone instead of yours. A pod stuck in `ImagePullBackOff` or `Pending` keeps the Job active and blocks every following run under `Forbid`.
How do I set a timezone on a Kubernetes CronJob? ▼
Use `spec.timeZone` with an IANA name such as `Europe/Paris`. It has been stable since Kubernetes 1.27. Putting `CRON_TZ=` or `TZ=` inside `schedule` is rejected by the API server, and `timeZone: Local` is rejected too. Without `timeZone`, the schedule is read in the kube-controller-manager's local time, which is usually UTC.
What happens when a Kubernetes CronJob misses more than 100 schedules? ▼
The docs say the CronJob does not start the Job and logs `too many missed start times. Set or decrease .spec.startingDeadlineSeconds or check clock skew`. In the current controller code, crossing 100 emits a `TooManyMissedTimes` warning event and still creates one Job for the most recent missed time. Either way, only one catch-up run happens: the other missed runs are gone.
Does a pod stuck in Pending count toward backoffLimit? ▼
No. `backoffLimit` counts failed pods, or container restarts with `restartPolicy: OnFailure`. A pod that never starts, because the image cannot be pulled or no node fits its requests, never fails, so the Job stays active until `activeDeadlineSeconds` fails it with `DeadlineExceeded`. Without that field it stays active indefinitely.
Can I monitor Kubernetes CronJobs with Prometheus? ▼
Yes, kube-state-metrics exposes `kube_cronjob_status_last_successful_time`, `kube_cronjob_next_schedule_time` and `kube_job_status_failed`, and you can alert on a last success that is too old. Those alerts depend on Prometheus and Alertmanager running inside the same cluster, so an external heartbeat is still the check that fires when the cluster itself is the problem.
How do I test a Kubernetes CronJob without waiting for the schedule? ▼
Run `kubectl create job backup-test-1 --from=cronjob/nightly-backup`. It creates a Job from the CronJob's template right away, with the annotation `cronjob.kubernetes.io/instantiate: manual`, and its pings reach the same healthcheck as the scheduled runs.




