To monitor a database backup, make the backup script check its own output (exit code, file size, a table you know must be there) and ping a heartbeat URL only when all of it passed. If that ping does not arrive on schedule, you get an alert. A backup that failed, wrote an empty file, hung, or never started all look the same from the outside: no success ping.
I'm Léo, I build Hyperping. This guide covers the ways pg_dump, mysqldump, restic and borg backups fail without telling anyone, the checks I put in backup scripts, and how I wire them to Hyperping healthchecks. I ran every script below on my Mac against PostgreSQL 18.6, MariaDB 13.0.2, restic 0.19.1 and borg 1.4.5, with a local listener in place of Hyperping, and broke each one on purpose.
Key takeaways
pg_dump wrongdb | gzip > backup.sql.gzexits 0 and leaves a 20 byte file. Addset -o pipefail, or don't pipe.pg_dumpof a database with no tables exits 0 with an 845 byte file. Check a minimum size and a table that must be present.- A dump killed halfway leaves a large file behind: 46 MB from
pg_dump, 8.4 MB frommysqldump. Write to.partialand rename only after the checks. - A mysqldump file without its
-- Dump completed onlast line is truncated. - Use
restic backup --stdin-from-commandandborg create --content-from-command, not a pipe: they fail when the dump command fails. - A backup you never restored is a guess. Run a weekly restore test with its own healthcheck.
How database backups fail silently
Backups are a cron job with a twist: when they break, nothing visible breaks with them. The app keeps working until the day you need the file.
GitLab's January 2017 outage is the reference case. Their postmortem explains that "the backup procedure was using pg_dump 9.2, while our database is running PostgreSQL 9.6", that "the S3 bucket was empty", and that the failure emails never arrived: "DMARC was not enabled for the cronjob emails, resulting in them being rejected by the receiver." The live incident notes describe backups "producing files only a few bytes in size." They lost six hours of data.
Here are the failure modes I reproduced.
A pipe returns the exit code of gzip
The Bash manual says it plainly: "The exit status of a pipeline is the exit status of the last command in the pipeline, unless the pipefail option is enabled." gzip happily compresses nothing:
$ pg_dump -d nosuchdb | gzip > backup.sql.gz; echo "exit=$?"
pg_dump: error: connection to server ... FATAL: database "nosuchdb" does not exist
exit=0
$ stat -c %s backup.sql.gz
20With set -o pipefail, the same line exits 1. Cron sees 0 without it, so a && curl ping after it would report success every night.
An empty database dumps fine
pg_dump does not know which database you meant. Point it at a fresh database with the same name (a new server after a failover, a wrong host in .pgpass) and you get a valid, tiny archive:
$ pg_dump -Fc -d empty -f empty.dump; echo "exit=$? size=$(stat -c %s empty.dump)"
exit=0 size=845Nothing failed, so nothing alerts. Only a size threshold or a content check catches it.
A dump killed halfway leaves a big file
When I terminated the pg_dump session in the middle of a 3 million row table, pg_dump exited 1 and left a 46 MB custom-format file. A mysqldump killed with kill -9 left an 8.4 MB file ending in the middle of an INSERT. Both look like backups in ls -l. The mysqldump reference even documents it for --result-file: the file is "created and its previous contents overwritten, even if an error occurs while generating the dump."
If your retention script keeps "the last 7 files", a week of partial files replaces your last good one.
restic and borg back up whatever they receive
The restic docs warn that with --stdin, "if mysqldump fails to connect to the MySQL database, the restic backup will nevertheless succeed in creating an empty backup" (restic backup docs). On restic 0.19.1, the piped version exited 3 ("at least one source file could not be read") and still saved a 0 byte snapshot. Exit code 3 is easy to treat as a warning and move on.
borg's docs say the same about piping into borg create -: it "can end up with truncated output being backed up" (borg create).
The job never runs
The cron file was not copied to the new server, the timer was never enabled, the disk is full so the script dies on its first write, the server was off at 3:30. None of these produce an error anyone reads. This is the case only an external check can see: something that expects a signal and notices when it does not come.
How to check backups with native tools
These commands tell you what happened, after the fact, when you think to run them:
# PostgreSQL custom-format dump: list the archive's contents
pg_restore --list app-2026-10-08-0330.dump | grep "TABLE DATA"
# mysqldump / mariadb-dump: a complete file ends with "-- Dump completed on <date>"
tail -n 1 app-2026-10-08-0330.sql
# restic: latest snapshot, its size, and a repository check
restic snapshots --latest 1
restic check --read-data-subset=5%
# borg: latest archive and its size
borg list --last 1
borg info --last 1
# cron output, if the job logs to a file
tail -n 50 /var/log/postgresql/backup.logpg_restore --list reads the table of contents without restoring anything (pg_restore docs). restic check verifies the repository's structure, and --read-data-subset also downloads and verifies a share of the data each time (restic docs). They are worth running. None of them alert anyone.
How to monitor database backups with Hyperping
A Hyperping healthcheck gives the backup a secret URL. If the success ping does not arrive by the scheduled time plus a grace period, Hyperping opens an incident and alerts you, and it resolves on the next successful ping. Healthchecks are included on every plan, Free included.
The trick with backups is where the success ping goes: after the checks, as the very last line of the script.
1. Create a healthcheck with the backup's schedule
In Hyperping, open Healthchecks, click Create healthcheck, and pick Cron. Use the same expression and the same timezone as the job. For a backup at 3:30 every night, Paris time, that is 30 3 * * * with Europe/Paris (the every day page of the cron generator explains each field).
I avoid 2:00 to 3:00 for nightly jobs in timezones with daylight saving time: that hour does not exist one night in spring, and cron implementations differ on what they do with it.
2. Write a backup script that checks its own output
This is the PostgreSQL version I use. It dumps in custom format to a .partial file, checks it, renames it, and pings success last.
#!/usr/bin/env bash
# /usr/local/bin/backup-postgres.sh
set -euo pipefail
: "${HC_URL:?set HC_URL}" # https://hc.hyperping.io/tok_...
DB="${DB:-app}"
BACKUP_DIR="${BACKUP_DIR:-/var/backups/postgres}"
MIN_BYTES="${MIN_BYTES:-1000000}" # the smallest dump you would believe
CHECK_TABLE="${CHECK_TABLE:-orders}" # a table that must be in every backup
ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }
ping /start
file="$BACKUP_DIR/$DB-$(date +%F-%H%M).dump"
pg_dump --no-password --format=custom --dbname="$DB" --file="$file.partial"
size=$(stat -c %s "$file.partial")
if [ "$size" -lt "$MIN_BYTES" ]; then
echo "Backup too small: $size bytes (minimum $MIN_BYTES)" >&2
exit 1
fi
toc=$(pg_restore --list "$file.partial")
if ! grep -q " TABLE DATA public $CHECK_TABLE " <<< "$toc"; then
echo "Backup has no data for table $CHECK_TABLE" >&2
exit 1
fi
mv "$file.partial" "$file"
find "$BACKUP_DIR" -name "$DB-*.dump" -mtime +14 -delete
ping ""A few details that matter:
set -euo pipefailmakes any failed command, unset variable or failed pipe stage stop the script before the success ping.--no-passwordmakespg_dumpfail instead of waiting for a password prompt that cron can never answer. Credentials go in~/.pgpassof the user running the job.- The
pingfunction never fails the script (|| true): an unreachable Hyperping must not turn a good backup into a failed one. With--retry 3, a short network blip does not cost you a false alert either. - I read the table of contents into a variable before grepping it.
pg_restore --list | grep -qunderpipefailcan fail on a good dump:grep -qexits at the first match andpg_restoregets SIGPIPE. - Set
MIN_BYTESfrom real numbers: look at a week of dump sizes and take about half of the smallest one. SetCHECK_TABLEto a table that always has rows.
Here is what each run sent, with MIN_BYTES=100000:
| Run | Exit | Pings received |
|---|---|---|
| Normal dump, 480 KB | 0 | /start, then success |
| Database with no tables (845 bytes) | 1, "Backup too small" | /start only |
| Wrong database that happens to be big | 1, "no data for table orders" | /start only |
| Database does not exist | 1, from pg_dump |
/start only |
The failed runs leave a .partial file behind, never a file that looks like a backup.
MySQL and MariaDB
mysqldump (or mariadb-dump, its current name in MariaDB) writes a completion trailer when comments are on, which is the default (MySQL reference, under --dump-date). The script checks it before compressing:
#!/usr/bin/env bash
# /usr/local/bin/backup-mysql.sh
set -euo pipefail
: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
BACKUP_DIR="${BACKUP_DIR:-/var/backups/mysql}"
MIN_BYTES="${MIN_BYTES:-1000000}"
ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }
ping /start
file="$BACKUP_DIR/$DB-$(date +%F-%H%M).sql"
mysqldump --defaults-extra-file=/etc/mysql/backup.cnf \
--single-transaction --quick --routines --triggers "$DB" > "$file.partial"
if ! tail -n 1 "$file.partial" | grep -q '^-- Dump completed'; then
echo "Dump has no completion trailer, it is truncated" >&2
exit 1
fi
size=$(stat -c %s "$file.partial")
if [ "$size" -lt "$MIN_BYTES" ]; then
echo "Backup too small: $size bytes (minimum $MIN_BYTES)" >&2
exit 1
fi
gzip "$file.partial" && mv "$file.partial.gz" "$file.gz"
find "$BACKUP_DIR" -name "$DB-*.sql.gz" -mtime +14 -delete
ping ""/etc/mysql/backup.cnf holds the credentials of a read-only backup user, mode 0600, so the password never shows up in ps:
[client]
user=backup
password=your-password--single-transaction gives a consistent snapshot of InnoDB tables without locking them, and --quick streams rows instead of buffering whole tables in memory. In my test, a wrong password exited 2 and left a 0 byte file, and the normal run ended with -- Dump completed on 2026-10-08 22:29:28.
restic
restic 0.17 added --stdin-from-command: restic runs the dump itself, and "a non-zero exit code from the command causes restic to cancel the backup" with no snapshot (restic docs). The size check then reads the snapshot summary:
#!/usr/bin/env bash
# /usr/local/bin/backup-restic.sh
# RESTIC_REPOSITORY and RESTIC_PASSWORD_FILE come from the environment.
set -euo pipefail
: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
MIN_BYTES="${MIN_BYTES:-1000000}"
ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }
ping /start
restic backup --quiet --tag postgres --stdin-filename "$DB.sql" \
--stdin-from-command -- pg_dump --no-password --dbname="$DB"
bytes=$(restic snapshots --tag postgres --latest 1 --json | jq '.[-1].summary.total_bytes_processed')
if [ "$bytes" -lt "$MIN_BYTES" ]; then
echo "Snapshot too small: $bytes bytes (minimum $MIN_BYTES)" >&2
exit 1
fi
ping ""With a database that does not exist, restic printed Fatal: unable to save snapshot: ... command failed: exit status 1, exited 1, and created no snapshot. The healthcheck got /start and nothing else.
borg
borg has had the same feature since 1.2: with --content-from-command, "borg is guaranteed to fail without creating an archive should the command fail" (borg create).
#!/usr/bin/env bash
# /usr/local/bin/backup-borg.sh
# BORG_REPO and BORG_PASSCOMMAND come from the environment.
set -euo pipefail
: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
MIN_BYTES="${MIN_BYTES:-1000000}"
ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }
ping /start
borg create --content-from-command --stdin-name "$DB.sql" \
"::$DB-{now:%Y-%m-%dT%H:%M}" -- pg_dump --no-password --dbname="$DB"
bytes=$(borg info --json --last 1 | jq '.archives[-1].stats.original_size')
if [ "$bytes" -lt "$MIN_BYTES" ]; then
echo "Archive too small: $bytes bytes (minimum $MIN_BYTES)" >&2
exit 1
fi
ping ""On borg 1.4.5, a failing pg_dump gave Command Error: Command 'pg_dump' exited with status 1 and exit code 2, with no archive. Keep in mind that borg exits 1 for warnings, such as a file that changed while it was read. set -e treats that as a failure, which is the safe default for a database dump.
3. Schedule it with the ping URL in the environment
With cron, a file in /etc/cron.d can set variables for its jobs:
# /etc/cron.d/backup-postgres
HC_URL=https://hc.hyperping.io/tok_your_backup_token
30 3 * * * postgres /usr/local/bin/backup-postgres.sh >> /var/log/postgresql/backup.log 2>&1Cron reads the time in the server's timezone, so the healthcheck must use the same one (run timedatectl to check it). Make the file readable by root only: the URL is the only credential of that healthcheck.
If you prefer systemd timers, my systemd timers guide shows the same job with EnvironmentFile= and a start timeout. For backups that run as Kubernetes CronJobs, see monitoring Kubernetes CronJobs. If your backups are scheduled inside PostgreSQL itself, monitoring pg_cron covers that case.
4. Add a weekly restore test with its own healthcheck
A dump that passes every check above can still fail to restore: a missing extension, a role the dump references, a format your new PostgreSQL major version reads differently. The only proof is a restore. I run this one every Sunday at 5:00, with a second healthcheck in cron mode, 0 5 * * 0 with Europe/Paris (every week runs on Sunday at midnight, change the hour).
#!/usr/bin/env bash
# /usr/local/bin/restore-test.sh
set -euo pipefail
: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
BACKUP_DIR="${BACKUP_DIR:-/var/backups/postgres}"
SCRATCH=restore_check
ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }
ping /start
latest=$(ls -1t "$BACKUP_DIR/$DB"-*.dump | head -n 1)
dropdb --if-exists "$SCRATCH"
createdb --template=template0 "$SCRATCH"
trap 'dropdb --if-exists "$SCRATCH"' EXIT
pg_restore --no-owner --exit-on-error --single-transaction --dbname="$SCRATCH" "$latest"
# The restored copy must have the tables and recent rows you expect.
rows=$(psql -Atc "SELECT count(*) FROM orders WHERE created_at > now() - interval '2 days'" "$SCRATCH")
if [ "$rows" -eq 0 ]; then
echo "Restored backup $latest has no orders from the last 2 days" >&2
exit 1
fi
ping ""--exit-on-error stops at the first error, where pg_restore would otherwise "continue and display a count of errors at the end", and --single-transaction means "either all the commands complete successfully, or no changes are applied" (pg_restore docs). Creating the scratch database from template0 keeps it empty, as the docs recommend.
The recent-rows query is what makes this test worth having. When I set every order's created_at to ten days ago and ran a new backup, the restore itself succeeded, and the script failed with "has no orders from the last 2 days". That is the stale-backup case: a dump that restores fine but comes from a replica that stopped replicating, or a database nobody writes to anymore. Run it on a separate machine when you can, since that also tests that the backup can leave the server.
5. Size the grace period and route the alerts
In cron mode, Hyperping expects the success ping by the scheduled time plus the grace period. Size it to the slowest normal run plus a margin: if the dump takes 12 to 20 minutes, 45 minutes of grace is fine. Because the script pings /start, each run's duration shows up in the healthcheck's ping history. After a week you know the real range, and a dump that suddenly takes 30 seconds instead of 15 minutes is worth a look even when it passes.
Alerts go to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so if a missed backup should page the person on call, send the alerts to PagerDuty or Opsgenie.
Test it once by hand
Run the script once as the user cron uses (sudo -u postgres env HC_URL=... /usr/local/bin/backup-postgres.sh) and check that both pings show up in the healthcheck. Then break it on purpose: set DB=nosuchdb, or MIN_BYTES above the real size, run it again, and wait for the alert after the grace period. Put it back and run it: the incident closes.
The same heartbeat pattern works for other schedulers: systemd timers, Node.js jobs, Rails scheduled jobs, and your monitoring stack itself with a Prometheus Alertmanager dead man's switch. If the backups run on a server you manage, my disk space guide covers the alert that usually fires before a backup dies on a full disk.
FAQ
How do I know if my database backup failed? ▼
Make the backup script prove its own result, then ping a heartbeat URL only when every check passed. Check the exit code of the dump tool (with `set -o pipefail` if you pipe it), the size of the file, and that a table you care about is in it. If the success ping does not arrive by the scheduled time plus a grace period, you get an alert, whether the script failed, hung, or never started.
Can pg_dump succeed and still produce an empty backup? ▼
Yes. On PostgreSQL 18.6, `pg_dump -Fc` of a database with no tables exits 0 and writes an 845 byte file. And `pg_dump wrongdb | gzip > backup.sql.gz` exits 0 with a 20 byte file, because a pipeline returns the exit code of its last command unless `set -o pipefail` is on. A minimum size and a check for a known table catch both.
How do I check that a mysqldump file is complete? ▼
With comments on (the default), mysqldump and mariadb-dump end the file with a line that starts with `-- Dump completed on`. A dump that was killed or cut off does not have it. Check the last line with `tail -n 1 dump.sql` before compressing the file, and treat a missing trailer as a failure.
Should I pipe pg_dump into restic backup --stdin? ▼
No. Use `restic backup --stdin-from-command -- pg_dump ...` (restic 0.17 and later). restic then runs the command itself, and a non-zero exit code cancels the backup without creating a snapshot. With a pipe, restic only sees an empty input: restic 0.19.1 exited 3 and still saved a 0 byte snapshot in my test. borg has the same option, `--content-from-command`.
How often should I test restoring a database backup? ▼
Automate it and run it at least weekly: restore the latest dump into a scratch database, check that a recent row is there, drop the scratch database. Give the restore test its own healthcheck, so a restore that breaks alerts you even while the nightly backups still look fine.
What grace period should a backup healthcheck have? ▼
Longer than the slowest normal run plus a margin. If the nightly dump takes 12 to 20 minutes, a 45 minute grace period is reasonable. Ping `/start` at the beginning of the script, and the healthcheck's ping history shows how long each run took, so you can tighten it after a week.




