To monitor a scheduled job in Rails, ping a heartbeat URL after the job's work succeeds (an around_perform callback does it for every Active Job), and alert when that ping does not arrive on schedule. Solid Queue, GoodJob and whenever each have their own way of skipping a run without telling anyone, and a missing ping catches all of them.
I'm Léo, I build Hyperping. This guide covers how Solid Queue recurring tasks, GoodJob cron and whenever fail silently, what Rails gives you to see it, and the concern I use with Hyperping healthchecks. I tested everything below on Ruby 4.0.7 and Rails 8.1.4, with Solid Queue 1.7.0 on SQLite, GoodJob 4.21.1 on PostgreSQL 18.6, and whenever 1.1.3, with a local listener in place of Hyperping.
Key takeaways
- Solid Queue recurring tasks are enqueued by its scheduler process. If
bin/jobsis down at 2:00, the 2:00 run is skipped and never caught up. - GoodJob skips missed cron runs too, unless you set
cron_graceful_restart_period. With it, every missed run in the window is enqueued at once. retry_onhandles the exception outside the perform callbacks, so anaround_performping afterblock.callonly fires on success.- whenever writes a crontab: no timezone option, and commands run in a login shell (
/bin/bash -l -c) with that user's Ruby. - Solid Queue has no automatic retries. A failed job sits in
solid_queue_failed_executionsuntil someone looks.
How Rails scheduled jobs fail silently
Solid Queue: the scheduler is a process you have to run
Rails 8 apps use Solid Queue by default. Recurring tasks live in config/recurring.yml, and the README says "the scheduler manages recurring tasks, enqueuing jobs for them when they're due." The scheduler is one of the processes bin/jobs starts. On a single server, Rails 8's default config/puma.rb can run it inside Puma instead:
# Run the Solid Queue supervisor inside of Puma for single-server deployments.
plugin :solid_queue if ENV["SOLID_QUEUE_IN_PUMA"]Kamal's generated deploy.yml sets SOLID_QUEUE_IN_PUMA: true. Move to a setup with separate web and job servers, forget to start bin/jobs, and recurring tasks stop. Nothing errors.
Missed runs are not caught up. I stopped bin/jobs for about 16 seconds on a 5 second schedule. The solid_queue_recurring_executions table shows the gap:
["20:47:15", "20:47:20", "20:47:25", "20:47:30", "20:47:35", "20:47:40", "20:48:00", "20:48:05", "20:48:10"]20:47:45, 20:47:50 and 20:47:55 never ran. On a daily task, a deploy or a crash at the scheduled minute costs you that day's run.
Solid Queue does not retry
The README is explicit: "Solid Queue doesn't include any automatic retry mechanism, it relies on Active Job for this." A recurring job that raises without retry_on gets a row in solid_queue_failed_executions and stays there "until manually discarded or re-enqueued". The next day's run is enqueued as usual, and fails the same way.
A job whose process was killed (out of memory, SIGKILL during a deploy) is marked failed with SolidQueue::Processes::ProcessPrunedError, and the README notes that "Active Job's retry_on and rescue_from have no effect on these errors." No callback of yours runs in that case.
Timezones depend on the gem and its version
- Solid Queue: a schedule can carry a zone (
0 2 * * * Europe/Paris). Without one, it usesconfig.time_zonesince version 1.5.0 (July 2026). Before that, the system's local time. - GoodJob: cron strings are parsed by Fugit, which accepts a zone after the expression. Without one, I read the GoodJob source as using the process's timezone, not
config.time_zone, so I always write the zone. - whenever: no timezone at all. Cron uses the server's.
On a server in UTC, 0 2 * * * written by someone in Paris runs at 3:00 or 4:00 their time depending on the season.
whenever: a crontab, a login shell, and no one reading the mail
whenever turns config/schedule.rb into crontab lines. Each one runs as /bin/bash -l -c '...', so the job uses the Ruby and gems that the cron user's login shell sets up. When I first ran a generated line on my Mac, the login shell picked the system Ruby 2.6 and the job died with a Bundler error before loading Rails. On a server, the same thing happens after a Ruby upgrade that only touched your interactive shell.
rails runner does exit with status 1 when the code raises (I checked: bin/rails runner -e production 'raise "boom"'; echo $? prints 1). Cron mails that output to MAILTO, if a mail transfer agent is installed and someone reads that inbox.
GoodJob cron: off by default, catch-up off by default
GoodJob runs on PostgreSQL only. Its cron needs config.good_job.enable_cron = true, which "defaults to false", in the process that should enqueue the jobs. Running it in several processes is safe: GoodJob uses unique indexes so each run is enqueued once.
cron_graceful_restart_period decides what happens to runs missed while GoodJob was down, and it "defaults to nil (disabled)". In my test, without it, a 16 second stop on a 5 second schedule skipped the missed runs. With GRACE=60, the restart enqueued all four missed runs (20:50:15 to 20:50:30) and ran them in the same second. For a daily job, a graceful period of a few minutes recovers a run missed during a deploy, which is what you want. For a job every minute, it means a burst after each restart.
How to check Rails scheduled jobs natively
# Solid Queue: recurring tasks, their recent runs, and failures
SolidQueue::RecurringTask.all.map { |t| [t.key, t.schedule] }
SolidQueue::RecurringExecution.where(task_key: "daily_report").order(run_at: :desc).limit(5).pluck(:run_at)
SolidQueue::FailedExecution.includes(:job).map { |f| [f.job.class_name, f.error["exception_class"]] }
# GoodJob: cron runs and their errors
GoodJob::Job.where(cron_key: "daily_report").order(cron_at: :desc).limit(5).pluck(:cron_at, :finished_at, :error)# whenever: what is installed in the crontab, and what schedule.rb would generate
crontab -l
bundle exec wheneverMission Control Jobs gives you a dashboard over Solid Queue, and GoodJob ships its own. Both show you failed jobs well. Neither shows a run that was never enqueued, and none of them sends an alert.
How to monitor Rails scheduled jobs with Hyperping
A Hyperping healthcheck gives each job a secret URL. If the success ping does not arrive by the scheduled time plus a grace period, Hyperping opens an incident and alerts you, and it resolves on the next successful ping. A scheduler that was down, a job that raised, a process killed mid-run: all of them end with no ping. Healthchecks are included on every plan, Free included.
1. Create a healthcheck with the job's schedule
In Hyperping, open Healthchecks, click Create healthcheck, and pick Cron. Copy the job's expression and the timezone the scheduler actually uses: 0 2 * * * with Europe/Paris for the examples below (the every day page of the cron generator explains each field). For Solid Queue's natural-language schedules, translate them: every day at 2am is 0 2 * * *, every hour is 0 * * * *. For every 15 minutes, use simple mode with the same interval.
2. Add a heartbeat concern to the job
# app/jobs/concerns/heartbeat.rb
require "net/http"
module Heartbeat
extend ActiveSupport::Concern
class_methods do
# heartbeat ENV["HC_URL_DAILY_REPORT"]
def heartbeat(url)
around_perform do |job, block|
Heartbeat.ping(url, "/start") if job.executions == 1
block.call
Heartbeat.ping(url) # skipped when perform raises
end
end
end
def self.ping(url, path = "")
return if url.blank?
Net::HTTP.post(URI("#{url}#{path}"), "")
rescue StandardError => e
Rails.logger.warn("Heartbeat ping failed: #{e.message}")
end
end# app/jobs/daily_report_job.rb
class DailyReportJob < ApplicationJob
include Heartbeat
queue_as :default
retry_on Net::OpenTimeout, wait: :polynomially_longer, attempts: 3
heartbeat ENV["HC_URL_DAILY_REPORT"]
def perform
DailyReport.deliver_all
end
endHow it behaves:
- When
performraises, the exception leaves the callback before the second ping. Withretry_on, the exception is rescued after the callbacks ran, so the failed attempt sends no success ping, and the retry that succeeds does. /startis sent on the first execution only (job.executionsis 1), so the duration Hyperping records covers the retries.pingrescues its own errors: an unreachable Hyperping must not fail a job that worked.
Here is what the listener received from that job on Solid Queue, on a 5 second schedule, with a job that raised Net::OpenTimeout on its second run and ArgumentError on its fourth:
| Run | What happened | Pings |
|---|---|---|
| 1 | success | /start, success |
| 2 | Net::OpenTimeout, retried by retry_on 4 seconds later, success |
/start, then success from the retry |
| 3 | ArgumentError, no retry_on for it |
/start only, job in solid_queue_failed_executions |
| 4 | success | /start, success |
GoodJob gave the same sequence. The user agent in the healthcheck's last pings is Ruby.
3. Schedule it with Solid Queue, GoodJob or whenever
Solid Queue
# config/recurring.yml
production:
daily_report:
class: DailyReportJob
schedule: "0 2 * * * Europe/Paris"Make sure the scheduler runs in production: bin/jobs as its own process (a systemd service, a Kamal role, a Procfile line), or Puma with SOLID_QUEUE_IN_PUMA=true on a single server. bin/jobs --skip-recurring and SOLID_QUEUE_SKIP_RECURRING=true turn the scheduler off, which you want on all but one kind of process if you run several.
GoodJob
# config/initializers/good_job.rb
Rails.application.configure do
config.active_job.queue_adapter = :good_job
config.good_job.enable_cron = true
config.good_job.cron_graceful_restart_period = 5.minutes
config.good_job.cron = {
daily_report: {
cron: "0 2 * * * Europe/Paris",
class: "DailyReportJob",
description: "Daily report emails",
},
}
endcron_graceful_restart_period set to your usual deploy time re-enqueues a run missed during a restart. enable_cron must be true in the process that runs good_job start (or in the web process, in async mode).
whenever
The ping goes in the crontab line, chained with && after rails runner, which exits 1 when the code raises. A custom job type keeps schedule.rb readable:
# config/schedule.rb
set :output, "log/cron.log"
env "HC_URL_DAILY_REPORT", ENV.fetch("HC_URL_DAILY_REPORT", "")
job_type :monitored_runner,
"curl -fsS -m 10 -o /dev/null $:hc/start; " \
"cd :path && :bundle_command :runner_command -e :environment ':task' :output && " \
"curl -fsS -m 10 --retry 3 -o /dev/null $:hc"
every 1.day, at: "2:00 am" do
monitored_runner "DailyReport.deliver_all", hc: "HC_URL_DAILY_REPORT"
endbundle exec whenever prints what it would install:
HC_URL_DAILY_REPORT=https://hc.hyperping.io/tok_your_report_token
0 2 * * * /bin/bash -l -c 'curl -fsS -m 10 -o /dev/null $HC_URL_DAILY_REPORT/start; cd /var/www/app && bundle exec bin/rails runner -e production '\''DailyReport.deliver_all'\'' >> log/cron.log 2>&1 && curl -fsS -m 10 --retry 3 -o /dev/null $HC_URL_DAILY_REPORT'Run HC_URL_DAILY_REPORT=... bundle exec whenever --update-crontab on the server (most deploy setups call it for you). When I ran that line by hand, the successful run sent /start and the success ping, and the run that raised sent /start only. The time is the server's: use the same timezone for the healthcheck, or convert.
If you call perform_now on an Active Job from the runner, the heartbeat concern pings too: give it a different healthcheck than the crontab line, or ping from one place only.
4. Size the grace period and route the alerts
In cron mode, Hyperping expects the success ping by the scheduled time plus the grace period. Cover the job's normal duration plus the retry_on waits: :polynomially_longer waits about 3, then 18 seconds plus jitter between attempts, so retries add little, but a job that also waits for a free worker on a busy queue can start minutes late. The healthcheck's ping history shows the time between /start and the success ping for each run.
Alerts go to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so send them to PagerDuty or Opsgenie when a missed job should page the person on call.
Test it once by hand
In a console on the server, run DailyReportJob.perform_now and check that /start and the success ping appear in the healthcheck. Then stop the scheduler (bin/jobs, or the GoodJob process) across one scheduled run, and wait for the alert after the grace period. Start it again: the next successful run closes the incident.
If you run Sidekiq instead of Solid Queue or GoodJob, my Sidekiq scheduled jobs guide covers sidekiq-cron, sidekiq-scheduler and Enterprise periodic jobs. The same heartbeat works for Node.js cron jobs, database backups and the Laravel scheduler. For the web side of the same app, see the Ruby on Rails health check endpoint guide.
FAQ
Why is my Solid Queue recurring task not running? ▼
Recurring tasks are enqueued by Solid Queue's scheduler process, which `bin/jobs` starts (or Puma, with `SOLID_QUEUE_IN_PUMA`). If that process is not running, or runs with `--skip-recurring`, nothing is enqueued. Also check that the task sits under the right environment key in `config/recurring.yml` (`production:`), and the timezone: since Solid Queue 1.5, a schedule without a zone uses `config.time_zone`.
Does Solid Queue catch up recurring tasks missed while it was down? ▼
No. On Solid Queue 1.7.0, I stopped `bin/jobs` for about 16 seconds on a 5 second schedule: the three runs due during the downtime never ran, and the scheduler resumed at the next slot. A deploy or crash at 2:00 skips a 2:00 task. GoodJob can re-enqueue missed runs with `cron_graceful_restart_period`, off by default.
Does after_perform run when a job fails or is retried with retry_on? ▼
No. `retry_on` rescues the exception outside the perform callbacks, so `after_perform`, and the code after `block.call` in an `around_perform`, are skipped for the failed attempt. They run on the attempt that succeeds. That makes `around_perform` a good place for a success-only heartbeat ping.
How do I set a timezone for whenever cron jobs? ▼
whenever has no timezone option: it writes a crontab, and cron uses the server's timezone. `env 'CRON_TZ', 'Europe/Paris'` in schedule.rb writes a CRON_TZ line, which cronie honors but Debian's cron does not. The simplest fix is to write the times in the server's timezone and use the same timezone for the healthcheck.
How do I see failed Solid Queue jobs? ▼
Solid Queue has no automatic retries: a job that raises without `retry_on` gets a row in `solid_queue_failed_executions` and stays there until you retry or discard it. Mission Control Jobs gives you a dashboard over them. Neither tells you about a recurring task that was never enqueued.




