To monitor a Node.js cron job, ping a heartbeat URL on the line after the job's work, and alert when that ping does not arrive on schedule. With node-cron, BullMQ or Agenda, a run can be skipped by a restart, swallowed by the scheduler's error handling, or retried until it succeeds, and none of that reaches you unless something expects a signal from each successful run.

I'm Léo, I build Hyperping. This guide covers how node-cron 4.6, BullMQ 6.3 (job schedulers) and Agenda 6.2 fail without telling you, what each one exposes natively, and the code I use with Hyperping healthchecks. I ran every snippet below on Node.js 20 against a local Redis 8.1, with a listener in place of Hyperping, and made each job fail on purpose. Use Node.js 24 in production: Node.js 20 reached end of life in April 2026.

Key takeaways

  • node-cron keeps the schedule in memory. A process that is down or restarting at 2:00 skips the 2:00 run, and v4 does not replay it.
  • An exception in a node-cron task is caught and logged, the process keeps running, and the next run happens as usual. No crash, no alert.
  • BullMQ's worker failed event fires on every failed attempt, including ones that succeed on retry. Alert on the missing success ping instead.
  • BullMQ 6 removed queue.add(..., { repeat }). Recurring jobs are job schedulers: queue.upsertJobScheduler().
  • Never put the success ping in finally: it reports failed runs as successful.

How Node.js scheduled jobs fail silently

node-cron dies with your process

node-cron is a timer inside your Node.js process. Its README says it "does not persist state to a database", and the v4 migration guide adds that "v4 does not replay missed runs automatically". So:

  • A deploy, crash loop or autoscaling event at the scheduled minute skips that run.
  • Two replicas of the same app run the job twice. Zero replicas run it zero times.
  • The schedule only exists in the processes that happened to call cron.schedule().

node-cron swallows errors

When the task throws or its promise rejects, node-cron catches it, logs it through its logger, emits execution:failed, and goes back to idle. I made a job throw on its second run:

[NODE-CRON] [ERROR] boom on run 2 Error: boom on run 2
event execution:failed boom on run 2
job ok run 3

The process kept running and the next run happened on time. That is the right behavior for a scheduler, and it means a job can fail every night with a line in a log nobody reads.

When I blocked the event loop for 5 seconds on a 2 second schedule, node-cron emitted execution:missed and ran one late execution. A CPU-heavy request handler in the same process can delay or skip your jobs.

The timezone is the process's

Without the timezone option, node-cron reads the pattern in the process's local time, which is UTC in most containers and on most cloud VMs. '0 2 * * *' written by someone in Paris runs at 3:00 or 4:00 their time depending on the season. Set it explicitly:

const task = cron.schedule('0 2 * * *', run, { timezone: 'Europe/Paris' });
console.log(task.getNextRun()); // 2026-10-09T00:00:00.000Z, that is 02:00 in Paris

BullMQ takes tz and Agenda takes timezone, with the same default.

BullMQ runs nothing without a worker, and collapses missed runs

BullMQ 6 removed the old repeatable jobs API: recurring jobs are now job schedulers, created with queue.upsertJobScheduler(). The docs explain that "there will be always one job associated to the scheduler in the 'Delayed' status" and that "the scheduler will only generate new jobs when the last job begins processing."

When I stopped the worker for 12 seconds on a 3 second schedule, the queue held a single delayed job. On restart, that job ran once, late, and the scheduler picked up the normal rhythm from the next slot. The three runs in between never happened, and nothing recorded that they were missed. In production, a worker deployment scaled to zero or crash looping on a bad environment variable looks exactly like this.

BullMQ also needs Redis with maxmemory-policy noeviction: the production guide calls it "the only setting that guarantees the correct behavior of the queues". A Redis that evicts keys under memory pressure can evict your scheduler.

BullMQ's failed event is not a failure alert

With attempts: 3, a job that fails once and succeeds on the second attempt is a success. The worker's failed event still fires for the first attempt. In my test:

daily-report attempt 1 failed: SMTP timeout
report sent 3

An alert wired to failed pages someone for a job that worked. Retries also stretch a run: with exponential backoff, the third attempt can start minutes after the scheduled time.

Agenda keeps the schedule in a database, with locks

Agenda 6 is a TypeScript rewrite with pluggable storage: MongoDB, PostgreSQL or Redis, each a separate package (@agendajs/mongo-backend, @agendajs/postgres-backend, @agendajs/redis-backend). The schedule survives restarts. In my tests on the Redis backend:

  • After a 10 second stop on a 3 second schedule, Agenda ran the job once on restart, then every 3 seconds. Missed runs collapse into one catch-up run.
  • A process killed in the middle of a run left the job locked. The interrupted run was never retried, and the job ran again at its next scheduled time once the lock had expired.

lockLifetime defaults to 10 minutes. A job that runs longer than that can be picked up by another worker while it is still running, so set it above your job's longest run.

How to check Node.js jobs natively

// node-cron: status, last run and next run of a task
console.log(task.getStatus(), task.lastRun(), task.getNextRun());

// BullMQ: schedulers, next run, and what is waiting or failed
console.log(await queue.getJobSchedulers());
console.log(await queue.getJobCounts('delayed', 'waiting', 'active', 'failed'));

// Agenda: stored jobs with their last run, next run and failure reason
console.log(await agenda.queryJobs({ name: 'daily report' }));

For BullMQ, Bull Board and Taskforce.sh give you a UI over the same data. All of these answer questions when you ask them. None of them notice a run that did not happen.

How to monitor Node.js cron jobs with Hyperping

A Hyperping healthcheck gives each job a secret URL. If the success ping does not arrive by the scheduled time plus a grace period, Hyperping opens an incident and alerts you, and it resolves on the next successful ping. It does not matter whether the process was down, the worker scaled to zero or the job threw: the ping is missing either way. Healthchecks are included on every plan, Free included.

1. Create a healthcheck with the job's schedule

In Hyperping, open Healthchecks, click Create healthcheck, and pick Cron. Use the job's pattern and timezone: 0 2 * * * with Europe/Paris for the examples below (the every day page of the cron generator explains each field). For BullMQ's every or Agenda's '15 minutes', use simple mode with the same interval, for example every 15 minutes.

Five-field patterns only: node-cron and BullMQ also accept a leading seconds field, which a healthcheck does not need for a daily job.

2. Add a ping helper that never throws

// heartbeat.js
const HC_URL = process.env.HC_URL; // https://hc.hyperping.io/tok_...

export async function ping(path = '') {
  if (!HC_URL) return;
  try {
    await fetch(HC_URL + path, { method: 'POST', signal: AbortSignal.timeout(10_000) });
  } catch (err) {
    console.error(`Heartbeat ping failed: ${err.message}`);
  }
}

fetch is global and stable in every supported Node.js version. The timeout keeps a slow network from holding the job, and the catch keeps an unreachable Hyperping from turning a good run into a failed one. Use one environment variable per job (HC_URL_DAILY_REPORT, and so on) when you monitor several.

3. Ping from node-cron, BullMQ or Agenda on success only

Each example pings /start before the work, which lets Hyperping record the run's duration, and the success URL on the line after the work. If sendDailyReport() throws, the success ping never runs.

node-cron

// cron.js
import cron from 'node-cron';
import { ping } from './heartbeat.js';
import { sendDailyReport } from './jobs.js';

const task = cron.schedule('0 2 * * *', async () => {
  await ping('/start');
  await sendDailyReport(); // throws on failure, so the success ping is skipped
  await ping();
}, { name: 'daily-report', timezone: 'Europe/Paris', noOverlap: true });

task.on('execution:failed', (ctx) => {
  console.error(`daily-report failed: ${ctx.execution?.error?.message}`);
});

noOverlap: true skips a run instead of starting a second one while the first is still going. If your app runs several replicas, run the scheduler in one of them only, or move the job to BullMQ or Agenda.

BullMQ job scheduler

Create the scheduler once, at deploy time or at boot. upsertJobScheduler is idempotent: calling it again with the same id updates the scheduler instead of adding a second one.

// queue.js
import { Queue } from 'bullmq';

const connection = { host: process.env.REDIS_HOST ?? '127.0.0.1', port: 6379 };
const queue = new Queue('reports', { connection });

await queue.upsertJobScheduler(
  'daily-report',
  { pattern: '0 2 * * *', tz: 'Europe/Paris' },
  { name: 'daily-report', opts: { attempts: 3, backoff: { type: 'exponential', delay: 60_000 } } },
);
await queue.close();
// worker.js
import { Worker } from 'bullmq';
import { ping } from './heartbeat.js';
import { sendDailyReport } from './jobs.js';

const connection = { host: process.env.REDIS_HOST ?? '127.0.0.1', port: 6379 };

const worker = new Worker('reports', async (job) => {
  if (job.name !== 'daily-report') return;
  if (job.attemptsMade === 0) await ping('/start');
  await sendDailyReport();
  await ping();
}, { connection });

worker.on('failed', (job, err) => {
  const final = job && job.attemptsMade >= (job.opts.attempts ?? 1);
  console.error(`${job?.name} attempt ${job?.attemptsMade} failed${final ? ' (no retries left)' : ''}: ${err.message}`);
});

Install ioredis next to bullmq (npm install bullmq ioredis): BullMQ 6 no longer installs it for you. /start is sent on the first attempt only, so the duration Hyperping records covers the retries too. A job that succeeds on its second attempt sends one success ping, and the healthcheck stays up.

Agenda

// agenda.js
import { Agenda } from 'agenda';
import { RedisBackend } from '@agendajs/redis-backend';
import { ping } from './heartbeat.js';
import { sendDailyReport } from './jobs.js';

const agenda = new Agenda({
  backend: new RedisBackend({ connectionString: process.env.REDIS_URL ?? 'redis://127.0.0.1:6379' }),
});

agenda.define('daily report', async () => {
  await ping('/start');
  await sendDailyReport();
  await ping();
}, { lockLifetime: 30 * 60 * 1000 });

agenda.on('fail:daily report', (err) => console.error(`daily report failed: ${err.message}`));

await agenda.start();
await agenda.every('0 2 * * *', 'daily report', undefined, { timezone: 'Europe/Paris' });

Swap RedisBackend for MongoBackend or PostgresBackend if that is where your data lives; the job code does not change. Agenda 6 is ESM only.

Here is what the three snippets sent to my listener, on a 3 second schedule with a job that fails on its second run:

Scheduler Run 1 Run 2 (throws) Run 3
node-cron /start, success /start only /start, success
BullMQ, attempts: 3 /start, success /start, then success on the retry /start, success
Agenda /start, success /start only /start, success

4. Keep the ping out of finally

This version looks tidy and defeats the purpose:

// Don't do this
cron.schedule('0 2 * * *', async () => {
  try {
    await sendDailyReport();
  } finally {
    await ping(); // runs when sendDailyReport() throws, too
  }
}, { timezone: 'Europe/Paris' });

The healthcheck receives a ping after every run, including failed ones, so it only alerts when the process is dead. To be alerted the moment a run fails rather than at its deadline, call the healthcheck's /fail URL from the catch block: Hyperping opens the incident at once. Keep finally for releasing locks, closing connections and deleting temporary files, and put the ping after the work.

5. Size the grace period and route the alerts

In cron mode, Hyperping expects the success ping by the scheduled time plus the grace period. Cover the job's normal duration, plus the retries and backoff delays if you use them. With attempts: 3 and an exponential backoff starting at 1 minute, the last attempt starts about 3 minutes after the first one fails, so a 5 minute job wants at least 30 minutes of grace. The healthcheck's ping history shows the time between /start and success for each run, so you can adjust after a week.

Alerts go to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so send them to PagerDuty or Opsgenie when a missed job should page the person on call.

Test it once by hand

Set the schedule a few minutes ahead (and the healthcheck to match), deploy, and check that /start and the success ping show up in the healthcheck's last pings with the user agent node. Then make the job throw, or stop the process, and wait for the alert after the grace period. Restore it: the next successful run closes the incident.

The same pattern applies to other runtimes: Rails scheduled jobs for Solid Queue, whenever and GoodJob, the Laravel scheduler, and Celery beat for Python. If your Node.js jobs run on a platform scheduler instead, see Vercel Cron Jobs and Kubernetes CronJobs. For the HTTP side of the same app, my guide to adding a health check endpoint with Express covers uptime monitoring. If a job calls an LLM, monitoring scheduled AI agents adds an output check before the ping.