To monitor a database backup, make the backup script check its own output (exit code, file size, a table you know must be there) and ping a heartbeat URL only when all of it passed. If that ping does not arrive on schedule, you get an alert. A backup that failed, wrote an empty file, hung, or never started all look the same from the outside: no success ping.

I'm Léo, I build Hyperping. This guide covers the ways pg_dump, mysqldump, restic and borg backups fail without telling anyone, the checks I put in backup scripts, and how I wire them to Hyperping healthchecks. I ran every script below on my Mac against PostgreSQL 18.6, MariaDB 13.0.2, restic 0.19.1 and borg 1.4.5, with a local listener in place of Hyperping, and broke each one on purpose.

Key takeaways

  • pg_dump wrongdb | gzip > backup.sql.gz exits 0 and leaves a 20 byte file. Add set -o pipefail, or don't pipe.
  • pg_dump of a database with no tables exits 0 with an 845 byte file. Check a minimum size and a table that must be present.
  • A dump killed halfway leaves a large file behind: 46 MB from pg_dump, 8.4 MB from mysqldump. Write to .partial and rename only after the checks.
  • A mysqldump file without its -- Dump completed on last line is truncated.
  • Use restic backup --stdin-from-command and borg create --content-from-command, not a pipe: they fail when the dump command fails.
  • A backup you never restored is a guess. Run a weekly restore test with its own healthcheck.

How database backups fail silently

Backups are a cron job with a twist: when they break, nothing visible breaks with them. The app keeps working until the day you need the file.

GitLab's January 2017 outage is the reference case. Their postmortem explains that "the backup procedure was using pg_dump 9.2, while our database is running PostgreSQL 9.6", that "the S3 bucket was empty", and that the failure emails never arrived: "DMARC was not enabled for the cronjob emails, resulting in them being rejected by the receiver." The live incident notes describe backups "producing files only a few bytes in size." They lost six hours of data.

Here are the failure modes I reproduced.

A pipe returns the exit code of gzip

The Bash manual says it plainly: "The exit status of a pipeline is the exit status of the last command in the pipeline, unless the pipefail option is enabled." gzip happily compresses nothing:

$ pg_dump -d nosuchdb | gzip > backup.sql.gz; echo "exit=$?"
pg_dump: error: connection to server ... FATAL:  database "nosuchdb" does not exist
exit=0
$ stat -c %s backup.sql.gz
20

With set -o pipefail, the same line exits 1. Cron sees 0 without it, so a && curl ping after it would report success every night.

An empty database dumps fine

pg_dump does not know which database you meant. Point it at a fresh database with the same name (a new server after a failover, a wrong host in .pgpass) and you get a valid, tiny archive:

$ pg_dump -Fc -d empty -f empty.dump; echo "exit=$? size=$(stat -c %s empty.dump)"
exit=0 size=845

Nothing failed, so nothing alerts. Only a size threshold or a content check catches it.

A dump killed halfway leaves a big file

When I terminated the pg_dump session in the middle of a 3 million row table, pg_dump exited 1 and left a 46 MB custom-format file. A mysqldump killed with kill -9 left an 8.4 MB file ending in the middle of an INSERT. Both look like backups in ls -l. The mysqldump reference even documents it for --result-file: the file is "created and its previous contents overwritten, even if an error occurs while generating the dump."

If your retention script keeps "the last 7 files", a week of partial files replaces your last good one.

restic and borg back up whatever they receive

The restic docs warn that with --stdin, "if mysqldump fails to connect to the MySQL database, the restic backup will nevertheless succeed in creating an empty backup" (restic backup docs). On restic 0.19.1, the piped version exited 3 ("at least one source file could not be read") and still saved a 0 byte snapshot. Exit code 3 is easy to treat as a warning and move on.

borg's docs say the same about piping into borg create -: it "can end up with truncated output being backed up" (borg create).

The job never runs

The cron file was not copied to the new server, the timer was never enabled, the disk is full so the script dies on its first write, the server was off at 3:30. None of these produce an error anyone reads. This is the case only an external check can see: something that expects a signal and notices when it does not come.

How to check backups with native tools

These commands tell you what happened, after the fact, when you think to run them:

# PostgreSQL custom-format dump: list the archive's contents
pg_restore --list app-2026-10-08-0330.dump | grep "TABLE DATA"

# mysqldump / mariadb-dump: a complete file ends with "-- Dump completed on <date>"
tail -n 1 app-2026-10-08-0330.sql

# restic: latest snapshot, its size, and a repository check
restic snapshots --latest 1
restic check --read-data-subset=5%

# borg: latest archive and its size
borg list --last 1
borg info --last 1

# cron output, if the job logs to a file
tail -n 50 /var/log/postgresql/backup.log

pg_restore --list reads the table of contents without restoring anything (pg_restore docs). restic check verifies the repository's structure, and --read-data-subset also downloads and verifies a share of the data each time (restic docs). They are worth running. None of them alert anyone.

How to monitor database backups with Hyperping

A Hyperping healthcheck gives the backup a secret URL. If the success ping does not arrive by the scheduled time plus a grace period, Hyperping opens an incident and alerts you, and it resolves on the next successful ping. Healthchecks are included on every plan, Free included.

The trick with backups is where the success ping goes: after the checks, as the very last line of the script.

1. Create a healthcheck with the backup's schedule

In Hyperping, open Healthchecks, click Create healthcheck, and pick Cron. Use the same expression and the same timezone as the job. For a backup at 3:30 every night, Paris time, that is 30 3 * * * with Europe/Paris (the every day page of the cron generator explains each field).

I avoid 2:00 to 3:00 for nightly jobs in timezones with daylight saving time: that hour does not exist one night in spring, and cron implementations differ on what they do with it.

2. Write a backup script that checks its own output

This is the PostgreSQL version I use. It dumps in custom format to a .partial file, checks it, renames it, and pings success last.

#!/usr/bin/env bash
# /usr/local/bin/backup-postgres.sh
set -euo pipefail

: "${HC_URL:?set HC_URL}"            # https://hc.hyperping.io/tok_...
DB="${DB:-app}"
BACKUP_DIR="${BACKUP_DIR:-/var/backups/postgres}"
MIN_BYTES="${MIN_BYTES:-1000000}"    # the smallest dump you would believe
CHECK_TABLE="${CHECK_TABLE:-orders}" # a table that must be in every backup

ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }

ping /start

file="$BACKUP_DIR/$DB-$(date +%F-%H%M).dump"
pg_dump --no-password --format=custom --dbname="$DB" --file="$file.partial"

size=$(stat -c %s "$file.partial")
if [ "$size" -lt "$MIN_BYTES" ]; then
  echo "Backup too small: $size bytes (minimum $MIN_BYTES)" >&2
  exit 1
fi

toc=$(pg_restore --list "$file.partial")
if ! grep -q " TABLE DATA public $CHECK_TABLE " <<< "$toc"; then
  echo "Backup has no data for table $CHECK_TABLE" >&2
  exit 1
fi

mv "$file.partial" "$file"
find "$BACKUP_DIR" -name "$DB-*.dump" -mtime +14 -delete

ping ""

A few details that matter:

  • set -euo pipefail makes any failed command, unset variable or failed pipe stage stop the script before the success ping.
  • --no-password makes pg_dump fail instead of waiting for a password prompt that cron can never answer. Credentials go in ~/.pgpass of the user running the job.
  • The ping function never fails the script (|| true): an unreachable Hyperping must not turn a good backup into a failed one. With --retry 3, a short network blip does not cost you a false alert either.
  • I read the table of contents into a variable before grepping it. pg_restore --list | grep -q under pipefail can fail on a good dump: grep -q exits at the first match and pg_restore gets SIGPIPE.
  • Set MIN_BYTES from real numbers: look at a week of dump sizes and take about half of the smallest one. Set CHECK_TABLE to a table that always has rows.

Here is what each run sent, with MIN_BYTES=100000:

Run Exit Pings received
Normal dump, 480 KB 0 /start, then success
Database with no tables (845 bytes) 1, "Backup too small" /start only
Wrong database that happens to be big 1, "no data for table orders" /start only
Database does not exist 1, from pg_dump /start only

The failed runs leave a .partial file behind, never a file that looks like a backup.

MySQL and MariaDB

mysqldump (or mariadb-dump, its current name in MariaDB) writes a completion trailer when comments are on, which is the default (MySQL reference, under --dump-date). The script checks it before compressing:

#!/usr/bin/env bash
# /usr/local/bin/backup-mysql.sh
set -euo pipefail

: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
BACKUP_DIR="${BACKUP_DIR:-/var/backups/mysql}"
MIN_BYTES="${MIN_BYTES:-1000000}"

ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }

ping /start

file="$BACKUP_DIR/$DB-$(date +%F-%H%M).sql"
mysqldump --defaults-extra-file=/etc/mysql/backup.cnf \
  --single-transaction --quick --routines --triggers "$DB" > "$file.partial"

if ! tail -n 1 "$file.partial" | grep -q '^-- Dump completed'; then
  echo "Dump has no completion trailer, it is truncated" >&2
  exit 1
fi

size=$(stat -c %s "$file.partial")
if [ "$size" -lt "$MIN_BYTES" ]; then
  echo "Backup too small: $size bytes (minimum $MIN_BYTES)" >&2
  exit 1
fi

gzip "$file.partial" && mv "$file.partial.gz" "$file.gz"
find "$BACKUP_DIR" -name "$DB-*.sql.gz" -mtime +14 -delete

ping ""

/etc/mysql/backup.cnf holds the credentials of a read-only backup user, mode 0600, so the password never shows up in ps:

[client]
user=backup
password=your-password

--single-transaction gives a consistent snapshot of InnoDB tables without locking them, and --quick streams rows instead of buffering whole tables in memory. In my test, a wrong password exited 2 and left a 0 byte file, and the normal run ended with -- Dump completed on 2026-10-08 22:29:28.

restic

restic 0.17 added --stdin-from-command: restic runs the dump itself, and "a non-zero exit code from the command causes restic to cancel the backup" with no snapshot (restic docs). The size check then reads the snapshot summary:

#!/usr/bin/env bash
# /usr/local/bin/backup-restic.sh
# RESTIC_REPOSITORY and RESTIC_PASSWORD_FILE come from the environment.
set -euo pipefail

: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
MIN_BYTES="${MIN_BYTES:-1000000}"

ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }

ping /start

restic backup --quiet --tag postgres --stdin-filename "$DB.sql" \
  --stdin-from-command -- pg_dump --no-password --dbname="$DB"

bytes=$(restic snapshots --tag postgres --latest 1 --json | jq '.[-1].summary.total_bytes_processed')
if [ "$bytes" -lt "$MIN_BYTES" ]; then
  echo "Snapshot too small: $bytes bytes (minimum $MIN_BYTES)" >&2
  exit 1
fi

ping ""

With a database that does not exist, restic printed Fatal: unable to save snapshot: ... command failed: exit status 1, exited 1, and created no snapshot. The healthcheck got /start and nothing else.

borg

borg has had the same feature since 1.2: with --content-from-command, "borg is guaranteed to fail without creating an archive should the command fail" (borg create).

#!/usr/bin/env bash
# /usr/local/bin/backup-borg.sh
# BORG_REPO and BORG_PASSCOMMAND come from the environment.
set -euo pipefail

: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
MIN_BYTES="${MIN_BYTES:-1000000}"

ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }

ping /start

borg create --content-from-command --stdin-name "$DB.sql" \
  "::$DB-{now:%Y-%m-%dT%H:%M}" -- pg_dump --no-password --dbname="$DB"

bytes=$(borg info --json --last 1 | jq '.archives[-1].stats.original_size')
if [ "$bytes" -lt "$MIN_BYTES" ]; then
  echo "Archive too small: $bytes bytes (minimum $MIN_BYTES)" >&2
  exit 1
fi

ping ""

On borg 1.4.5, a failing pg_dump gave Command Error: Command 'pg_dump' exited with status 1 and exit code 2, with no archive. Keep in mind that borg exits 1 for warnings, such as a file that changed while it was read. set -e treats that as a failure, which is the safe default for a database dump.

3. Schedule it with the ping URL in the environment

With cron, a file in /etc/cron.d can set variables for its jobs:

# /etc/cron.d/backup-postgres
HC_URL=https://hc.hyperping.io/tok_your_backup_token
30 3 * * * postgres /usr/local/bin/backup-postgres.sh >> /var/log/postgresql/backup.log 2>&1

Cron reads the time in the server's timezone, so the healthcheck must use the same one (run timedatectl to check it). Make the file readable by root only: the URL is the only credential of that healthcheck.

If you prefer systemd timers, my systemd timers guide shows the same job with EnvironmentFile= and a start timeout. For backups that run as Kubernetes CronJobs, see monitoring Kubernetes CronJobs. If your backups are scheduled inside PostgreSQL itself, monitoring pg_cron covers that case.

4. Add a weekly restore test with its own healthcheck

A dump that passes every check above can still fail to restore: a missing extension, a role the dump references, a format your new PostgreSQL major version reads differently. The only proof is a restore. I run this one every Sunday at 5:00, with a second healthcheck in cron mode, 0 5 * * 0 with Europe/Paris (every week runs on Sunday at midnight, change the hour).

#!/usr/bin/env bash
# /usr/local/bin/restore-test.sh
set -euo pipefail

: "${HC_URL:?set HC_URL}"
DB="${DB:-app}"
BACKUP_DIR="${BACKUP_DIR:-/var/backups/postgres}"
SCRATCH=restore_check

ping() { curl -fsS -m 10 --retry 3 -o /dev/null "${HC_URL}$1" || true; }

ping /start

latest=$(ls -1t "$BACKUP_DIR/$DB"-*.dump | head -n 1)

dropdb --if-exists "$SCRATCH"
createdb --template=template0 "$SCRATCH"
trap 'dropdb --if-exists "$SCRATCH"' EXIT
pg_restore --no-owner --exit-on-error --single-transaction --dbname="$SCRATCH" "$latest"

# The restored copy must have the tables and recent rows you expect.
rows=$(psql -Atc "SELECT count(*) FROM orders WHERE created_at > now() - interval '2 days'" "$SCRATCH")
if [ "$rows" -eq 0 ]; then
  echo "Restored backup $latest has no orders from the last 2 days" >&2
  exit 1
fi

ping ""

--exit-on-error stops at the first error, where pg_restore would otherwise "continue and display a count of errors at the end", and --single-transaction means "either all the commands complete successfully, or no changes are applied" (pg_restore docs). Creating the scratch database from template0 keeps it empty, as the docs recommend.

The recent-rows query is what makes this test worth having. When I set every order's created_at to ten days ago and ran a new backup, the restore itself succeeded, and the script failed with "has no orders from the last 2 days". That is the stale-backup case: a dump that restores fine but comes from a replica that stopped replicating, or a database nobody writes to anymore. Run it on a separate machine when you can, since that also tests that the backup can leave the server.

5. Size the grace period and route the alerts

In cron mode, Hyperping expects the success ping by the scheduled time plus the grace period. Size it to the slowest normal run plus a margin: if the dump takes 12 to 20 minutes, 45 minutes of grace is fine. Because the script pings /start, each run's duration shows up in the healthcheck's ping history. After a week you know the real range, and a dump that suddenly takes 30 seconds instead of 15 minutes is worth a look even when it passes.

Alerts go to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so if a missed backup should page the person on call, send the alerts to PagerDuty or Opsgenie.

Test it once by hand

Run the script once as the user cron uses (sudo -u postgres env HC_URL=... /usr/local/bin/backup-postgres.sh) and check that both pings show up in the healthcheck. Then break it on purpose: set DB=nosuchdb, or MIN_BYTES above the real size, run it again, and wait for the alert after the grace period. Put it back and run it: the incident closes.

The same heartbeat pattern works for other schedulers: systemd timers, Node.js jobs, Rails scheduled jobs, and your monitoring stack itself with a Prometheus Alertmanager dead man's switch. If the backups run on a server you manage, my disk space guide covers the alert that usually fires before a backup dies on a full disk.