To monitor a scheduled job in Rails, ping a heartbeat URL after the job's work succeeds (an around_perform callback does it for every Active Job), and alert when that ping does not arrive on schedule. Solid Queue, GoodJob and whenever each have their own way of skipping a run without telling anyone, and a missing ping catches all of them.

I'm Léo, I build Hyperping. This guide covers how Solid Queue recurring tasks, GoodJob cron and whenever fail silently, what Rails gives you to see it, and the concern I use with Hyperping healthchecks. I tested everything below on Ruby 4.0.7 and Rails 8.1.4, with Solid Queue 1.7.0 on SQLite, GoodJob 4.21.1 on PostgreSQL 18.6, and whenever 1.1.3, with a local listener in place of Hyperping.

Key takeaways

  • Solid Queue recurring tasks are enqueued by its scheduler process. If bin/jobs is down at 2:00, the 2:00 run is skipped and never caught up.
  • GoodJob skips missed cron runs too, unless you set cron_graceful_restart_period. With it, every missed run in the window is enqueued at once.
  • retry_on handles the exception outside the perform callbacks, so an around_perform ping after block.call only fires on success.
  • whenever writes a crontab: no timezone option, and commands run in a login shell (/bin/bash -l -c) with that user's Ruby.
  • Solid Queue has no automatic retries. A failed job sits in solid_queue_failed_executions until someone looks.

How Rails scheduled jobs fail silently

Solid Queue: the scheduler is a process you have to run

Rails 8 apps use Solid Queue by default. Recurring tasks live in config/recurring.yml, and the README says "the scheduler manages recurring tasks, enqueuing jobs for them when they're due." The scheduler is one of the processes bin/jobs starts. On a single server, Rails 8's default config/puma.rb can run it inside Puma instead:

# Run the Solid Queue supervisor inside of Puma for single-server deployments.
plugin :solid_queue if ENV["SOLID_QUEUE_IN_PUMA"]

Kamal's generated deploy.yml sets SOLID_QUEUE_IN_PUMA: true. Move to a setup with separate web and job servers, forget to start bin/jobs, and recurring tasks stop. Nothing errors.

Missed runs are not caught up. I stopped bin/jobs for about 16 seconds on a 5 second schedule. The solid_queue_recurring_executions table shows the gap:

["20:47:15", "20:47:20", "20:47:25", "20:47:30", "20:47:35", "20:47:40", "20:48:00", "20:48:05", "20:48:10"]

20:47:45, 20:47:50 and 20:47:55 never ran. On a daily task, a deploy or a crash at the scheduled minute costs you that day's run.

Solid Queue does not retry

The README is explicit: "Solid Queue doesn't include any automatic retry mechanism, it relies on Active Job for this." A recurring job that raises without retry_on gets a row in solid_queue_failed_executions and stays there "until manually discarded or re-enqueued". The next day's run is enqueued as usual, and fails the same way.

A job whose process was killed (out of memory, SIGKILL during a deploy) is marked failed with SolidQueue::Processes::ProcessPrunedError, and the README notes that "Active Job's retry_on and rescue_from have no effect on these errors." No callback of yours runs in that case.

Timezones depend on the gem and its version

  • Solid Queue: a schedule can carry a zone (0 2 * * * Europe/Paris). Without one, it uses config.time_zone since version 1.5.0 (July 2026). Before that, the system's local time.
  • GoodJob: cron strings are parsed by Fugit, which accepts a zone after the expression. Without one, I read the GoodJob source as using the process's timezone, not config.time_zone, so I always write the zone.
  • whenever: no timezone at all. Cron uses the server's.

On a server in UTC, 0 2 * * * written by someone in Paris runs at 3:00 or 4:00 their time depending on the season.

whenever: a crontab, a login shell, and no one reading the mail

whenever turns config/schedule.rb into crontab lines. Each one runs as /bin/bash -l -c '...', so the job uses the Ruby and gems that the cron user's login shell sets up. When I first ran a generated line on my Mac, the login shell picked the system Ruby 2.6 and the job died with a Bundler error before loading Rails. On a server, the same thing happens after a Ruby upgrade that only touched your interactive shell.

rails runner does exit with status 1 when the code raises (I checked: bin/rails runner -e production 'raise "boom"'; echo $? prints 1). Cron mails that output to MAILTO, if a mail transfer agent is installed and someone reads that inbox.

GoodJob cron: off by default, catch-up off by default

GoodJob runs on PostgreSQL only. Its cron needs config.good_job.enable_cron = true, which "defaults to false", in the process that should enqueue the jobs. Running it in several processes is safe: GoodJob uses unique indexes so each run is enqueued once.

cron_graceful_restart_period decides what happens to runs missed while GoodJob was down, and it "defaults to nil (disabled)". In my test, without it, a 16 second stop on a 5 second schedule skipped the missed runs. With GRACE=60, the restart enqueued all four missed runs (20:50:15 to 20:50:30) and ran them in the same second. For a daily job, a graceful period of a few minutes recovers a run missed during a deploy, which is what you want. For a job every minute, it means a burst after each restart.

How to check Rails scheduled jobs natively

# Solid Queue: recurring tasks, their recent runs, and failures
SolidQueue::RecurringTask.all.map { |t| [t.key, t.schedule] }
SolidQueue::RecurringExecution.where(task_key: "daily_report").order(run_at: :desc).limit(5).pluck(:run_at)
SolidQueue::FailedExecution.includes(:job).map { |f| [f.job.class_name, f.error["exception_class"]] }

# GoodJob: cron runs and their errors
GoodJob::Job.where(cron_key: "daily_report").order(cron_at: :desc).limit(5).pluck(:cron_at, :finished_at, :error)
# whenever: what is installed in the crontab, and what schedule.rb would generate
crontab -l
bundle exec whenever

Mission Control Jobs gives you a dashboard over Solid Queue, and GoodJob ships its own. Both show you failed jobs well. Neither shows a run that was never enqueued, and none of them sends an alert.

How to monitor Rails scheduled jobs with Hyperping

A Hyperping healthcheck gives each job a secret URL. If the success ping does not arrive by the scheduled time plus a grace period, Hyperping opens an incident and alerts you, and it resolves on the next successful ping. A scheduler that was down, a job that raised, a process killed mid-run: all of them end with no ping. Healthchecks are included on every plan, Free included.

1. Create a healthcheck with the job's schedule

In Hyperping, open Healthchecks, click Create healthcheck, and pick Cron. Copy the job's expression and the timezone the scheduler actually uses: 0 2 * * * with Europe/Paris for the examples below (the every day page of the cron generator explains each field). For Solid Queue's natural-language schedules, translate them: every day at 2am is 0 2 * * *, every hour is 0 * * * *. For every 15 minutes, use simple mode with the same interval.

2. Add a heartbeat concern to the job

# app/jobs/concerns/heartbeat.rb
require "net/http"

module Heartbeat
  extend ActiveSupport::Concern

  class_methods do
    # heartbeat ENV["HC_URL_DAILY_REPORT"]
    def heartbeat(url)
      around_perform do |job, block|
        Heartbeat.ping(url, "/start") if job.executions == 1
        block.call
        Heartbeat.ping(url) # skipped when perform raises
      end
    end
  end

  def self.ping(url, path = "")
    return if url.blank?
    Net::HTTP.post(URI("#{url}#{path}"), "")
  rescue StandardError => e
    Rails.logger.warn("Heartbeat ping failed: #{e.message}")
  end
end
# app/jobs/daily_report_job.rb
class DailyReportJob < ApplicationJob
  include Heartbeat

  queue_as :default
  retry_on Net::OpenTimeout, wait: :polynomially_longer, attempts: 3
  heartbeat ENV["HC_URL_DAILY_REPORT"]

  def perform
    DailyReport.deliver_all
  end
end

How it behaves:

  • When perform raises, the exception leaves the callback before the second ping. With retry_on, the exception is rescued after the callbacks ran, so the failed attempt sends no success ping, and the retry that succeeds does.
  • /start is sent on the first execution only (job.executions is 1), so the duration Hyperping records covers the retries.
  • ping rescues its own errors: an unreachable Hyperping must not fail a job that worked.

Here is what the listener received from that job on Solid Queue, on a 5 second schedule, with a job that raised Net::OpenTimeout on its second run and ArgumentError on its fourth:

Run What happened Pings
1 success /start, success
2 Net::OpenTimeout, retried by retry_on 4 seconds later, success /start, then success from the retry
3 ArgumentError, no retry_on for it /start only, job in solid_queue_failed_executions
4 success /start, success

GoodJob gave the same sequence. The user agent in the healthcheck's last pings is Ruby.

3. Schedule it with Solid Queue, GoodJob or whenever

Solid Queue

# config/recurring.yml
production:
  daily_report:
    class: DailyReportJob
    schedule: "0 2 * * * Europe/Paris"

Make sure the scheduler runs in production: bin/jobs as its own process (a systemd service, a Kamal role, a Procfile line), or Puma with SOLID_QUEUE_IN_PUMA=true on a single server. bin/jobs --skip-recurring and SOLID_QUEUE_SKIP_RECURRING=true turn the scheduler off, which you want on all but one kind of process if you run several.

GoodJob

# config/initializers/good_job.rb
Rails.application.configure do
  config.active_job.queue_adapter = :good_job
  config.good_job.enable_cron = true
  config.good_job.cron_graceful_restart_period = 5.minutes
  config.good_job.cron = {
    daily_report: {
      cron: "0 2 * * * Europe/Paris",
      class: "DailyReportJob",
      description: "Daily report emails",
    },
  }
end

cron_graceful_restart_period set to your usual deploy time re-enqueues a run missed during a restart. enable_cron must be true in the process that runs good_job start (or in the web process, in async mode).

whenever

The ping goes in the crontab line, chained with && after rails runner, which exits 1 when the code raises. A custom job type keeps schedule.rb readable:

# config/schedule.rb
set :output, "log/cron.log"
env "HC_URL_DAILY_REPORT", ENV.fetch("HC_URL_DAILY_REPORT", "")

job_type :monitored_runner,
  "curl -fsS -m 10 -o /dev/null $:hc/start; " \
  "cd :path && :bundle_command :runner_command -e :environment ':task' :output && " \
  "curl -fsS -m 10 --retry 3 -o /dev/null $:hc"

every 1.day, at: "2:00 am" do
  monitored_runner "DailyReport.deliver_all", hc: "HC_URL_DAILY_REPORT"
end

bundle exec whenever prints what it would install:

HC_URL_DAILY_REPORT=https://hc.hyperping.io/tok_your_report_token

0 2 * * * /bin/bash -l -c 'curl -fsS -m 10 -o /dev/null $HC_URL_DAILY_REPORT/start; cd /var/www/app && bundle exec bin/rails runner -e production '\''DailyReport.deliver_all'\'' >> log/cron.log 2>&1 && curl -fsS -m 10 --retry 3 -o /dev/null $HC_URL_DAILY_REPORT'

Run HC_URL_DAILY_REPORT=... bundle exec whenever --update-crontab on the server (most deploy setups call it for you). When I ran that line by hand, the successful run sent /start and the success ping, and the run that raised sent /start only. The time is the server's: use the same timezone for the healthcheck, or convert.

If you call perform_now on an Active Job from the runner, the heartbeat concern pings too: give it a different healthcheck than the crontab line, or ping from one place only.

4. Size the grace period and route the alerts

In cron mode, Hyperping expects the success ping by the scheduled time plus the grace period. Cover the job's normal duration plus the retry_on waits: :polynomially_longer waits about 3, then 18 seconds plus jitter between attempts, so retries add little, but a job that also waits for a free worker on a busy queue can start minutes late. The healthcheck's ping history shows the time between /start and the success ping for each run.

Alerts go to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so send them to PagerDuty or Opsgenie when a missed job should page the person on call.

Test it once by hand

In a console on the server, run DailyReportJob.perform_now and check that /start and the success ping appear in the healthcheck. Then stop the scheduler (bin/jobs, or the GoodJob process) across one scheduled run, and wait for the alert after the grace period. Start it again: the next successful run closes the incident.

If you run Sidekiq instead of Solid Queue or GoodJob, my Sidekiq scheduled jobs guide covers sidekiq-cron, sidekiq-scheduler and Enterprise periodic jobs. The same heartbeat works for Node.js cron jobs, database backups and the Laravel scheduler. For the web side of the same app, see the Ruby on Rails health check endpoint guide.