To set up a dead man's switch for Prometheus, route the always-firing Watchdog alert to an Alertmanager webhook that calls a heartbeat URL every few minutes, and alert when those calls stop. If Prometheus stops evaluating rules, Alertmanager crashes, or the network between Alertmanager and the outside world breaks, the pings stop and the heartbeat service tells you, through a channel that does not depend on the pipeline that just broke.

I'm Léo, I build Hyperping. This guide covers why an alerting stack cannot report its own failure, the configuration I use with Hyperping healthchecks for plain Alertmanager and for kube-prometheus-stack, and the timing details I measured. I ran Prometheus 3.15.0 and Alertmanager 0.34.1 (the current releases) on my Mac with a local listener in place of Hyperping, killed Prometheus mid-test, and checked how hc.hyperping.io answers Alertmanager's request.

Key takeaways

  • kube-prometheus-stack ships the Watchdog alert but routes it to the null receiver. Out of the box, it goes nowhere.
  • Alertmanager's webhook sends a POST with a JSON body and User-Agent: Alertmanager/0.34.1. Hyperping healthchecks accept it and ignore the body.
  • Set send_resolved: false. Otherwise Alertmanager sends one last "resolved" POST after Prometheus dies, and the heartbeat counts it as a ping.
  • repeat_interval only fires on a group_interval tick. With 1m and 20s, I measured pings every 60 to 80 seconds.
  • Use one healthcheck per Alertmanager or cluster. A shared URL keeps getting pings from the healthy clusters and hides the dead one.

Why an alerting pipeline fails silently

Prometheus alerting has three moving parts: Prometheus evaluates rules and sends firing alerts to Alertmanager, Alertmanager groups them and sends notifications to receivers (Slack, PagerDuty, email), and each receiver delivers to a person. Any of them can break, and the thing that should warn you is the thing that broke:

  • Prometheus is down, OOM-killed, or stuck on a full disk, so no rule is evaluated.
  • Prometheus cannot reach Alertmanager (wrong service name after an upgrade, network policy).
  • Alertmanager is down or crash-looping on a bad config.
  • A receiver stopped working: revoked Slack webhook, expired PagerDuty key, an SMTP relay that now rejects the mail.
  • The whole cluster is down, along with everything that would have reported it.

Prometheus has alerts for some of this, such as AlertmanagerFailedToSendAlerts, and they all travel through the same pipeline. The Watchdog runbook puts it simply: "If not firing then it should alert external systems that this alerting system is no longer working."

The Watchdog alert

kube-prometheus and kube-prometheus-stack include this rule in the general.rules group:

- alert: Watchdog
  expr: vector(1)
  labels:
    severity: none
  annotations:
    summary: An alert that should always be firing to certify that Alertmanager is working properly.

vector(1) always returns a value, so the alert always fires. Its description says it is "meant to ensure that the entire alerting pipeline is functional" and that "there are integrations with various notification mechanisms that send a notification when this alert is not firing." That is the dead man's switch.

If you run Prometheus without kube-prometheus, add the rule above to any rule file.

kube-prometheus-stack sends it to the null receiver

The chart's default values contain this route:

route:
  group_by: ['namespace']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 12h
  receiver: 'null'
  routes:
  - receiver: 'null'
    matchers:
      - alertname = "Watchdog"
receivers:
- name: 'null'

The Watchdog fires in every cluster installed with the chart, and its notifications are discarded. Many teams see it in the Alertmanager UI, assume it is wired somewhere, and move on.

How to check the pipeline with native tools

# Is Watchdog firing in Prometheus?
curl -s http://prometheus:9090/api/v1/alerts | jq '.data.alerts[] | select(.labels.alertname=="Watchdog") | .state'

# Did Alertmanager receive it? Add --silenced or --inhibited to see it if it is muted
amtool alert query alertname=Watchdog --alertmanager.url=http://alertmanager:9093

# Which receiver does it go to?
amtool config routes test --config.file=alertmanager.yml alertname=Watchdog

# Notification errors
kubectl logs -n monitoring -l app.kubernetes.io/name=alertmanager -c alertmanager | grep "Notify for alerts failed"

These tell you the state of the pipeline when you look. None of them can tell you anything when the cluster running them is down.

How to monitor Prometheus and Alertmanager with Hyperping

A Hyperping healthcheck is a secret URL that expects a ping on a schedule. Alertmanager pings it by sending the Watchdog notification to it. When the pings stop for longer than the period plus the grace period, Hyperping opens an incident and alerts you, from outside your infrastructure. Healthchecks are included on every plan, Free included.

1. Create a healthcheck in simple mode

In Hyperping, open Healthchecks, click Create healthcheck, and pick the simple mode: a ping expected every 5 minutes, with a 5 minute grace period. Cron mode is for jobs that run at fixed times; the Watchdog's rhythm comes from Alertmanager's timers, so an interval fits better.

Create one healthcheck per Alertmanager cluster, named after it (alertmanager prod-eu). If three clusters share one URL, a dead cluster is hidden by the pings of the two others.

2. Route Watchdog to a webhook receiver

For a plain Alertmanager, add a child route above your other routes:

# alertmanager.yml
route:
  receiver: default
  group_by: [alertname]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - matchers:
        - alertname = "Watchdog"
      receiver: hyperping-heartbeat
      group_wait: 0s
      group_interval: 1m
      repeat_interval: 5m

receivers:
  - name: default
    # your usual receivers
  - name: hyperping-heartbeat
    webhook_configs:
      - url_file: /etc/alertmanager/secrets/hyperping-heartbeat/url
        send_resolved: false

What each setting does here, from the Alertmanager configuration reference:

  • repeat_interval: 5m is the heartbeat rhythm. The docs say it "should be a multiple of the group_interval. If it's not, the repeat_interval is rounded up to the next multiple of the group_interval."
  • group_interval: 1m is how often Alertmanager checks whether a repeat is due. A repeat fires on the first tick after repeat_interval has passed, so it can land one tick late.
  • send_resolved: false stops the "resolved" notification that would count as a ping (more on that below).

I measured the first point with repeat_interval: 1m and group_interval: 20s: pings arrived 80, 60, 80, 80 and 80 seconds apart. With 5m and 1m, expect 5 to 6 minutes, which the 5 minute period plus 5 minute grace covers.

Hyperping limits each healthcheck to 10 pings per minute. A single Watchdog group sends one ping per repeat, so you are far from it, even with several Alertmanager replicas.

3. Keep the ping URL in a file

The URL is the only credential of the healthcheck. url_file (Alertmanager 0.26 and later) reads it from a file, so it stays out of the config you commit:

echo -n 'https://hc.hyperping.io/tok_your_watchdog_token' > /etc/alertmanager/secrets/hyperping-heartbeat/url

With kube-prometheus-stack, create a Secret and let the operator mount it. Secrets listed in alertmanagerSpec.secrets are mounted into /etc/alertmanager/secrets/<secret-name> in the Alertmanager container (Prometheus Operator API):

kubectl create secret generic hyperping-heartbeat -n monitoring \
  --from-literal=url='https://hc.hyperping.io/tok_your_watchdog_token'
# values.yaml for kube-prometheus-stack
alertmanager:
  alertmanagerSpec:
    secrets:
      - hyperping-heartbeat
  config:
    route:
      group_by: ['namespace']
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 12h
      receiver: 'null'
      routes:
        - receiver: hyperping-heartbeat
          matchers:
            - alertname = "Watchdog"
          group_wait: 0s
          group_interval: 1m
          repeat_interval: 5m
    receivers:
      - name: 'null'
      - name: hyperping-heartbeat
        webhook_configs:
          - url_file: /etc/alertmanager/secrets/hyperping-heartbeat/url
            send_resolved: false

Helm replaces lists instead of merging them, so this routes list replaces the chart's default one, including its Watchdog route to null. Add your other routes and receivers to the same lists.

I put this route in the main config rather than in an AlertmanagerConfig resource. By default, the operator adds a namespace matcher to every route that comes from an AlertmanagerConfig ("only process alerts that have a namespace label equal to the namespace of the object", per the operator's API), and vector(1) produces an alert without a namespace label, so such a route would not match it.

4. Check the route and the first pings

Before reloading Alertmanager, check the file and the route:

$ amtool check-config alertmanager.yml
Checking 'alertmanager.yml'  SUCCESS
$ amtool config routes test --config.file=alertmanager.yml --verify.receivers=hyperping-heartbeat alertname=Watchdog severity=none
hyperping-heartbeat

After the reload, open the healthcheck: the last pings should show POST requests with the user agent Alertmanager/0.34.1 (your version), and the healthcheck goes up.

Alertmanager's request is a POST with Content-Type: application/json and a JSON body ("version": "4", "status": "firing", the alerts). Hyperping's ping endpoint accepts HEAD, GET and POST, ignores the body, and answers 200, which counts as a successful notification. If the URL is wrong, it answers 401 and Alertmanager logs:

level=ERROR msg="Notify for alerts failed" ... err="hyperping-heartbeat/webhook[0]: notify retry canceled due to unrecoverable error after 1 attempts: unexpected status code 401: Token not valid"

The webhook only retries 5xx responses and network errors; a 4xx is final. Since the failed notification is never recorded as sent, Alertmanager tries again at the next group_interval tick, so a fixed URL starts pinging within a minute.

5. Send the alerts somewhere that does not depend on Alertmanager

When the Watchdog pings stop, the alert comes from Hyperping, not from your cluster, so it arrives even if the cluster is gone. Alerts go to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so to wake up whoever is on call, connect PagerDuty or Opsgenie to the project. Avoid relying only on a channel that your Alertmanager also uses through the same integration: if that integration is what broke, you want a second path.

What happens when Prometheus dies

Prometheus does not tell Alertmanager that it stopped. Each alert it sends carries an end time, and in the rule evaluation code that end time is four times the larger of the evaluation interval and the resend delay ("Allow for two Eval or Alertmanager send failures"). With the default resend delay of 1 minute, Alertmanager keeps the Watchdog firing for up to 4 minutes after the last evaluation.

Here is what I recorded after killing Prometheus at 22:13:42, with two webhooks on the same route, one with send_resolved: false and one with the default:

Time send_resolved: false default (true)
22:13:25 firing firing
22:13:42 Prometheus killed Prometheus killed
22:14:45 firing firing
22:16:05 firing firing
22:17:05 nothing resolved
after nothing nothing

With the default, the last ping is the "resolved" notification, one cycle later. With send_resolved: false, the pings simply stop. Either way, count on the end time: with a 5 minute period and 5 minute grace, the alert comes roughly 10 to 15 minutes after Prometheus stops, and 10 minutes after the last ping when Alertmanager itself stops. For a faster alert, use repeat_interval: 2m, group_interval: 1m, and a 3 minute period with 3 minutes of grace.

Two more things stop the pings and are worth knowing:

  • A silence that matches Watchdog. I added amtool silence add alertname=Watchdog, and the pings stopped until I expired it. A broad maintenance silence (one matching every alert) will set off the dead man's switch, which is correct but surprising.
  • Inhibition rules. An inhibit_rules entry whose target matches Watchdog mutes it too. amtool alert query --inhibited alertname=Watchdog lists it when that happens.

In a highly available Alertmanager cluster, replicas share a notification log and only one of them sends each notification; the others wait their turn ("wait_time = peer_position × peer_timeout", in the HA docs). During a network partition, you may get duplicate pings, which is harmless.

Test it once by hand

Silence the Watchdog for 20 minutes (amtool silence add alertname=Watchdog --duration=20m --comment="dead man's switch test") and wait for the Hyperping alert after the period plus the grace period. Expire the silence: the next notification is a ping, and the incident resolves.

The same heartbeat covers the jobs your cluster runs on a schedule: see monitoring Kubernetes CronJobs, and database backups for the job you least want to fail silently. Node.js and Rails apps have their own guides: Node.js cron jobs and Rails scheduled jobs. For alert routing on the uptime side, my notes on alert management cover what I page on and what I don't.