To monitor Celery beat, give each periodic task its own heartbeat URL and ping it from the worker when the task succeeds, with a task_success signal handler. If beat is down, the message sits in a queue no worker reads, or the task raises, the ping does not arrive and you get an alert. The trap is that beat only publishes messages: its log prints Sending due task on schedule whether or not anything ever runs the task.

I'm Léo, I build Hyperping, and this guide covers the ways Celery beat fails without an error, what Celery's own tools show you, and the setup I use with Hyperping healthchecks. I tested everything below with Celery 5.6.3 (the current stable release in October 2026), kombu 5.6.2 and django-celery-beat 2.9.0 against Redis, with a local HTTP listener standing in for the ping endpoint: on macOS with Python 3.14 and the threads pool, and on Ubuntu 22.04 with Python 3.10 and the default prefork pool.

Key takeaways

  • Beat publishes, workers execute. In my test, beat logged Sending due task every 10 seconds while 10 messages piled up in a queue no worker consumed, and celery status reported the worker online.
  • Run exactly one beat. Two beats send every task twice, and pointing them at the same celerybeat-schedule file does not stop it.
  • Ping from task_success, filtered on the names of your periodic tasks. It fires in the worker after the task returned, so a task that raises, or runs out of retries, never pings.
  • The django-celery-beat admin's last_run_at and total_run_count count messages sent: they went up with no worker running at all.
  • Monitor from outside: one healthcheck per periodic task in cron mode with the Celery timezone, plus a five minute heartbeat task that proves beat and a worker both work.

How Celery beat fails silently

Celery beat is a scheduler process. It sleeps until the next entry is due, publishes a message to the broker, and goes back to sleep. A worker, in another process and often on another machine, picks the message up and runs the task. Each link in that chain can break without anything turning red.

Beat is not running

Beat is its own process, started with celery -A proj beat. The docs allow embedding it in a worker with -B, which is "convenient if you'll never run more than one worker node, but it's not commonly used and for that reason isn't recommended for production use".

So beat is the process that gets forgotten: a Procfile or Compose file with a worker service and no beat service, a systemd unit that was never enabled on the new server, a Deployment scaled to zero during an incident. The workers stay up, the queues stay empty, and no error is logged anywhere, because nothing is running to log it.

Beat also needs a writable schedule file. I started it with -s pointing to a read-only directory, and it exited with code 1 on PermissionError: [Errno 13] Permission denied. In a container with a read-only root filesystem, beat crashes on start, and a supervisor that stops retrying leaves it dead. Pass -s /var/run/celery/celerybeat-schedule with a path that is writable.

Two beats send every task twice

The docs are direct about it: "You have to ensure only a single scheduler is running for a schedule at a time, otherwise you'd end up with duplicate tasks." Celery has no lock to enforce it.

The usual causes are worker -B on a worker scaled to three replicas, or a beat Deployment with two replicas or a rolling update where the old and new pods overlap. I ran two beats against the same schedule file. On Python 3.10, where shelve uses gdbm, the second beat logged this and carried on:

ERROR/MainProcess] Removing corrupted schedule file '/root/.../sched-shared': error(11, 'Resource temporarily unavailable')

It read the lock as corruption, deleted the file, created a new one, and both beats sent every heartbeat. On Python 3.14, where shelve defaults to SQLite, there was no message at all, and the tasks were still sent twice. On Kubernetes, run beat as a single replica with strategy: Recreate.

The schedule file is reset

The default PersistentScheduler stores the last run time of each entry in the celerybeat-schedule file. beat.py deletes and recreates it when it cannot be read ("Removing corrupted schedule file"), and clears it when the timezone or enable_utc changes. I switched timezone from Europe/Paris to UTC and got Reset: Timezone changed from 'Europe/Paris' to 'UTC'.

A cleared file has a side effect. A new entry starts with "last run" set to now, so a crontab task whose time passed while beat was down is never caught up. I checked crontab(minute=0, hour=2) with beat coming back at 02:10: with the file intact, the task was due immediately; with a fresh file, its next run was the following night. A beat container without a volume for the schedule file loses it on every restart.

django-celery-beat has its own rules

With django-celery-beat, the schedule lives in the database and beat runs with -S django. Its DatabaseScheduler checks the tables every 5 seconds (the banner shows maxinterval -> 5.00 seconds (5s)), so edits in the admin apply without a restart. A few behaviors stay quiet:

  • Unticking Enabled stops the task, and saving a disabled task resets its last_run_at to empty. Beat logs DatabaseScheduler: Schedule changed. and nothing else.
  • A periodic task with an Expires date disables itself once that date passes.
  • Each crontab schedule has its own timezone field. After changing Django's TIME_ZONE, the README says the schedule "will still be based on the old timezone" until you run PeriodicTask.objects.all().update(last_run_at=None) and PeriodicTasks.update_changed().
  • last_run_at and total_run_count are set by beat when it sends the message. I ran beat with no worker at all and watched both go up.

The timezone is not the one you think

crontab() entries are evaluated in the app's timezone: the timezone setting, "UTC" by default. In a Django project, Celery falls back to TIME_ZONE when CELERY_TIMEZONE is not set (I checked both cases), and with enable_utc = False and no timezone, it uses the server's local time.

Daylight saving time moves runs too. I stepped beat's is_due() check minute by minute across the spring change in Europe/Paris: on March 28, 2027, crontab(minute=30, hour=2) ran at 03:30, because 02:30 does not exist that day. Keep critical tasks out of the 02:00 to 03:00 window, or run Celery in UTC.

One syntax trap: minute defaults to "*", so crontab(hour=7) runs 60 times, every minute from 07:00 to 07:59. Write crontab(minute=0, hour=7).

Beat sent it, but no worker ran it

Beat's job ends when the message reaches the broker. I gave an entry "options": {"queue": "reports"} and ran no worker on that queue. For 100 seconds, beat logged Sending due task report-unrouted (tasks.report) every 10 seconds, celery status showed one node online, inspect active was empty, and LLEN reports in Redis reached 10. Nothing reported a problem.

The same happens when every worker is down, or when a backlog delays the queue for hours. Without an expiry, the queued copies all run in a burst when a worker comes back. With "options": {"expires": 20} on the entry, the worker I started on reports discarded the 10 stale messages with Discarding revoked task: tasks.report[...].

Beat reads beat_schedule once, at startup, and sends task names. If a deploy renames a task and restarts the workers but not beat, beat keeps sending the old name, and the worker logs Received unregistered task of type 'tasks.old_name' and drops the message. Restart beat on every deploy.

Late acks and the visibility timeout run a task twice

With acks_late=True, a worker acknowledges the message after the task finishes instead of before it starts. On Redis, a message that is not acknowledged within the visibility timeout "will be redelivered to another worker and executed". The default is one hour, so a nightly export that runs for 70 minutes with late acks gets delivered again while it is still running.

I reproduced it on Linux with a 10 second visibility timeout and a 60 second task. When a second worker started 20 seconds in, it received the same task ID and ran it in parallel. Both runs succeeded and both sent their pings. task_reject_on_worker_lost adds another path: a worker process killed mid-task puts the message back in the queue, and the task starts over. Set visibility_timeout above your longest task, and make periodic tasks safe to run twice.

Retries hide failures until they run out

autoretry_for and self.retry() turn a failure into a new execution of the same task. I ran a task that raised ConnectionError on its first two attempts with autoretry_for=(ConnectionError,), and another that called self.retry() with max_retries=2 and always failed:

22:08:53 GET /tok_flaky/start
22:08:56 GET /tok_flaky/start
22:08:59 GET /tok_flaky/start
22:08:59 GET /tok_flaky                  <- third attempt succeeded
22:08:53 GET /tok_doomed/start
22:08:56 GET /tok_doomed/start
22:08:59 GET /tok_doomed/start           <- gave up, no success ping

task_prerun fires on every attempt and task_success only on the one that returns, so a success ping is the right signal. Retries do push it later, though: self.retry() waits default_retry_delay, 3 minutes, between attempts by default, and retry_backoff=True can wait up to retry_backoff_max, 600 seconds, per attempt. The grace period has to cover that.

How to check Celery beat with its own tools

Celery gives you a good view of the workers:

# Workers that answer a ping (beat is not one of them)
celery -A proj status

# Task names each worker knows, to catch a renamed task
celery -A proj inspect registered

# Tasks running right now
celery -A proj inspect active

# Tasks a worker holds with an ETA or countdown, such as a pending retry
celery -A proj inspect scheduled

# Live task events; start the workers with -E first
celery -A proj events --dump

inspect scheduled does not list the beat schedule. It returned only my failing task's pending retry, with its eta. The beat schedule itself is visible in the beat log, one INFO line per message sent: Scheduler: Sending due task heartbeat (tasks.heartbeat). For history, Flower (celery -A proj flower) builds a task list from the same events, and with django-celery-beat the admin shows last_run_at and total_run_count per periodic task.

All of these report what happened. When beat is dead, it writes no log line, celery status still lists healthy workers, Flower shows an idle cluster, and the admin's last_run_at stops moving without anyone looking at it. Nothing alerts on something that did not happen. That is the case an external heartbeat covers.

How to monitor Celery beat with Hyperping

A Hyperping healthcheck is a secret URL that expects a request on a schedule. When the request does not arrive by the expected time plus a grace period, it opens an incident and alerts you, then resolves on the next successful ping. Healthchecks are included on every plan, Free included.

1. Create one healthcheck per periodic task

In Hyperping, open Healthchecks, click Create healthcheck, name it after the beat entry, and pick Cron for crontab() entries. Use the same expression, and the same timezone as the Celery timezone setting:

Celery schedule Healthcheck Generator page
crontab(minute="*/15") Cron */15 * * * * every 15 minutes
crontab(minute=0) Cron 0 * * * * every hour
crontab(minute=0, hour=0) Cron 0 0 * * * every day
crontab(minute=30, hour=2) Cron 30 2 * * *
crontab(minute=0, hour=8, day_of_week="mon-fri") Cron 0 8 * * 1-5
crontab(minute=0, hour=9, day_of_week="sun") Cron 0 9 * * 0
crontab(minute=0, hour=6, day_of_month=1) Cron 0 6 1 * *
timedelta(minutes=10) or 600.0 Simple, every 10 minutes

I checked each row by stepping Celery's own is_due() minute by minute and comparing the runs with the cron expression's, in Europe/Paris. They matched on ordinary days and across the October 2026 clock change. The one difference was the spring-forward day described above, where Celery ran the 02:30 task at 03:30. With django-celery-beat, use the timezone stored on that crontab schedule. In day_of_week, Celery counts like cron: 0 or "sun" is Sunday, 1 is Monday. It rejects 7 ("Valid range is 0-6"), which cron also accepts as Sunday.

One difference has no cron equivalent. When both day_of_month and day_of_week are set, Celery requires both: crontab(minute=0, hour=9, day_of_week="mon", day_of_month="1-7") ran only on the first Monday of each month in my test. Cron reads the same fields as "either", so 0 9 1-7 * 1 would expect a ping every Monday. For that schedule, run the task daily on days 1 to 7 with 0 9 1-7 * *, and let the task return early unless it is Monday.

Each healthcheck gives you a URL like https://hc.hyperping.io/tok_....

2. Store the ping URLs in environment variables

The pings are sent by the worker that runs the task, so the URLs belong in the workers' environment. Beat does not need them.

# Worker environment (.env, systemd EnvironmentFile, or a Kubernetes Secret)
HC_NIGHTLY_EXPORT_URL=https://hc.hyperping.io/tok_your_export_token
HC_SEND_INVOICES_URL=https://hc.hyperping.io/tok_your_invoices_token
HC_BEAT_HEARTBEAT_URL=https://hc.hyperping.io/tok_your_heartbeat_token

3. Ping from a task_success signal handler

task_success is "dispatched when a task succeeds", with the task object as sender. It runs in the worker process, after the result is stored, and only when the task returned without raising. task_prerun runs before each execution, which is where /start goes:

# proj/healthchecks.py
import logging
import os
import urllib.request

from celery.signals import task_prerun, task_success

logger = logging.getLogger(__name__)

# One healthcheck per periodic task, keyed by the task name beat sends
HEALTHCHECKS = {
    "proj.tasks.nightly_export": os.environ.get("HC_NIGHTLY_EXPORT_URL"),
    "proj.tasks.send_invoices": os.environ.get("HC_SEND_INVOICES_URL"),
    "proj.tasks.beat_heartbeat": os.environ.get("HC_BEAT_HEARTBEAT_URL"),
}


def ping(url, task_name):
    # Never let monitoring break or slow down the task
    if not url:
        return
    try:
        urllib.request.urlopen(url, timeout=5).close()
    except Exception as exc:
        logger.warning("Healthcheck ping for %s failed: %r", task_name, exc)


@task_prerun.connect
def healthcheck_start(sender=None, **kwargs):
    url = HEALTHCHECKS.get(sender.name)
    if url:
        ping(f"{url}/start", sender.name)


@task_success.connect
def healthcheck_success(sender=None, **kwargs):
    url = HEALTHCHECKS.get(sender.name)
    if url:
        ping(url, sender.name)

Import the module where your Celery app is created, so every worker connects the handlers at startup:

# proj/celery.py
import os

from celery import Celery

os.environ.setdefault("DJANGO_SETTINGS_MODULE", "proj.settings")

app = Celery("proj")
app.config_from_object("django.conf:settings", namespace="CELERY")
app.autodiscover_tasks()

from proj import healthchecks  # noqa: E402,F401  (connects the signal handlers)

A few choices in that code matter:

  • The handlers fire for every task the worker runs, including the thousands of on-demand ones. Filtering on sender.name keeps one ping per periodic run, and keeps you under the limit of 10 pings a minute per healthcheck.
  • Celery catches exceptions raised by signal handlers and logs Signal handler ... raised, so a failed ping cannot fail the task. A hanging ping would still hold the worker slot, hence the 5 second timeout. With the listener stopped, I got Healthcheck ping for tasks.heartbeat failed: URLError(ConnectionRefusedError(...)) and the task still succeeded.
  • Send the ping inline. Sending it as another Celery task would make the monitoring depend on the same broker and workers it is supposed to watch.
  • task_prerun fires again on each retry, so Hyperping measures the duration from the last attempt's /start. For a task that runs more often than once a minute, drop /start: a task every 10 seconds with both pings sends 12 requests a minute.
  • A manual .delay() of the same task also pings. If a task runs both on a schedule and on demand, schedule a thin wrapper task with its own name and monitor that.

If you prefer to keep it next to the work, ping at the end of the task body instead:

# proj/tasks.py
import os

from celery import shared_task

from proj.healthchecks import ping


@shared_task(bind=True, autoretry_for=(ConnectionError,), retry_backoff=True, max_retries=3)
def nightly_export(self):
    export_orders()
    ping(os.environ.get("HC_NIGHTLY_EXPORT_URL"), self.name)

The try/except in ping() matters even more here: an exception raised by the ping inside the task body would mark the export as failed, or with a broad autoretry_for, run it again.

4. Add a heartbeat task for beat and the workers

A weekly task only notices a dead beat a week later. Schedule a no-op task every five minutes:

# proj/tasks.py
from celery import shared_task


@shared_task
def beat_heartbeat():
    return "ok"
# proj/celery.py, after the app is created
from celery.schedules import crontab

app.conf.beat_schedule = {
    "beat-heartbeat": {
        "task": "proj.tasks.beat_heartbeat",
        "schedule": crontab(minute="*/5"),
        "options": {"expires": 240},
    },
    # ... your other periodic tasks
}

Its success ping proves the whole chain works: beat sent the message, the broker delivered it, and a worker ran it. Create its healthcheck in cron mode with */5 * * * * (or simple mode, every 5 minutes) and a 5 minute grace period. If beat dies, the default queue stops being consumed, or every worker is down, you know within about ten minutes. The expires option drops heartbeats that waited in the queue for more than four minutes, so a backlog cannot pass for a healthy cluster.

If your periodic tasks go to several queues, add one heartbeat per queue with "options": {"queue": "reports", "expires": 240} and its own healthcheck.

5. Size the grace period to queue wait and run time

In cron mode, Hyperping expects the success ping by the scheduled time plus the grace period. The ping is sent when the task ends, so the grace period has to cover the time the message waits in the queue, the run time, and the retry delays. A 20 minute export that fails fast on a connection error and retries up to three times, five minutes apart, needs at least 35 minutes, and I would give it 45. If a retry can follow a long partial run, add that run time too.

After a few runs, the healthcheck's ping history shows the duration between each /start and success ping. Set the grace period to roughly twice the longest one, plus the retry delays. The default is 10 minutes, and the minimum is 1.

6. Route the alerts

A missed ping goes to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so if a failed nightly export has to wake someone, send the alert to PagerDuty or Opsgenie and let their rotation decide who gets paged.

Test it before you trust it

Send the task by hand with celery -A proj call proj.tasks.nightly_export. The handler filters on the task name, not on who sent it, so the worker pings like it would for a scheduled run. Check the healthcheck's last pings: a STARTED entry, then an OK entry with Python-urllib/3.x as the user agent. Then make the task raise on staging and wait: no success ping arrives, and the alert should reach you once the grace period passes. Stop beat for fifteen minutes too, and the heartbeat healthcheck should go down while the workers look perfectly healthy.

If you run scheduled jobs outside Celery, the same pattern works for Sidekiq cron jobs, pg_cron, Kubernetes CronJobs and systemd timers. For a side by side of heartbeat tools, see the best cron job monitoring tools. Ruby and Node.js apps have their own guides: Rails scheduled jobs and Node.js cron jobs.