To set up a dead man's switch for Prometheus, route the always-firing Watchdog alert to an Alertmanager webhook that calls a heartbeat URL every few minutes, and alert when those calls stop. If Prometheus stops evaluating rules, Alertmanager crashes, or the network between Alertmanager and the outside world breaks, the pings stop and the heartbeat service tells you, through a channel that does not depend on the pipeline that just broke.
I'm Léo, I build Hyperping. This guide covers why an alerting stack cannot report its own failure, the configuration I use with Hyperping healthchecks for plain Alertmanager and for kube-prometheus-stack, and the timing details I measured. I ran Prometheus 3.15.0 and Alertmanager 0.34.1 (the current releases) on my Mac with a local listener in place of Hyperping, killed Prometheus mid-test, and checked how hc.hyperping.io answers Alertmanager's request.
Key takeaways
- kube-prometheus-stack ships the Watchdog alert but routes it to the
nullreceiver. Out of the box, it goes nowhere. - Alertmanager's webhook sends a POST with a JSON body and
User-Agent: Alertmanager/0.34.1. Hyperping healthchecks accept it and ignore the body. - Set
send_resolved: false. Otherwise Alertmanager sends one last "resolved" POST after Prometheus dies, and the heartbeat counts it as a ping. repeat_intervalonly fires on agroup_intervaltick. With 1m and 20s, I measured pings every 60 to 80 seconds.- Use one healthcheck per Alertmanager or cluster. A shared URL keeps getting pings from the healthy clusters and hides the dead one.
Why an alerting pipeline fails silently
Prometheus alerting has three moving parts: Prometheus evaluates rules and sends firing alerts to Alertmanager, Alertmanager groups them and sends notifications to receivers (Slack, PagerDuty, email), and each receiver delivers to a person. Any of them can break, and the thing that should warn you is the thing that broke:
- Prometheus is down, OOM-killed, or stuck on a full disk, so no rule is evaluated.
- Prometheus cannot reach Alertmanager (wrong service name after an upgrade, network policy).
- Alertmanager is down or crash-looping on a bad config.
- A receiver stopped working: revoked Slack webhook, expired PagerDuty key, an SMTP relay that now rejects the mail.
- The whole cluster is down, along with everything that would have reported it.
Prometheus has alerts for some of this, such as AlertmanagerFailedToSendAlerts, and they all travel through the same pipeline. The Watchdog runbook puts it simply: "If not firing then it should alert external systems that this alerting system is no longer working."
The Watchdog alert
kube-prometheus and kube-prometheus-stack include this rule in the general.rules group:
- alert: Watchdog
expr: vector(1)
labels:
severity: none
annotations:
summary: An alert that should always be firing to certify that Alertmanager is working properly.vector(1) always returns a value, so the alert always fires. Its description says it is "meant to ensure that the entire alerting pipeline is functional" and that "there are integrations with various notification mechanisms that send a notification when this alert is not firing." That is the dead man's switch.
If you run Prometheus without kube-prometheus, add the rule above to any rule file.
kube-prometheus-stack sends it to the null receiver
The chart's default values contain this route:
route:
group_by: ['namespace']
group_wait: 30s
group_interval: 5m
repeat_interval: 12h
receiver: 'null'
routes:
- receiver: 'null'
matchers:
- alertname = "Watchdog"
receivers:
- name: 'null'The Watchdog fires in every cluster installed with the chart, and its notifications are discarded. Many teams see it in the Alertmanager UI, assume it is wired somewhere, and move on.
How to check the pipeline with native tools
# Is Watchdog firing in Prometheus?
curl -s http://prometheus:9090/api/v1/alerts | jq '.data.alerts[] | select(.labels.alertname=="Watchdog") | .state'
# Did Alertmanager receive it? Add --silenced or --inhibited to see it if it is muted
amtool alert query alertname=Watchdog --alertmanager.url=http://alertmanager:9093
# Which receiver does it go to?
amtool config routes test --config.file=alertmanager.yml alertname=Watchdog
# Notification errors
kubectl logs -n monitoring -l app.kubernetes.io/name=alertmanager -c alertmanager | grep "Notify for alerts failed"These tell you the state of the pipeline when you look. None of them can tell you anything when the cluster running them is down.
How to monitor Prometheus and Alertmanager with Hyperping
A Hyperping healthcheck is a secret URL that expects a ping on a schedule. Alertmanager pings it by sending the Watchdog notification to it. When the pings stop for longer than the period plus the grace period, Hyperping opens an incident and alerts you, from outside your infrastructure. Healthchecks are included on every plan, Free included.
1. Create a healthcheck in simple mode
In Hyperping, open Healthchecks, click Create healthcheck, and pick the simple mode: a ping expected every 5 minutes, with a 5 minute grace period. Cron mode is for jobs that run at fixed times; the Watchdog's rhythm comes from Alertmanager's timers, so an interval fits better.
Create one healthcheck per Alertmanager cluster, named after it (alertmanager prod-eu). If three clusters share one URL, a dead cluster is hidden by the pings of the two others.
2. Route Watchdog to a webhook receiver
For a plain Alertmanager, add a child route above your other routes:
# alertmanager.yml
route:
receiver: default
group_by: [alertname]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- matchers:
- alertname = "Watchdog"
receiver: hyperping-heartbeat
group_wait: 0s
group_interval: 1m
repeat_interval: 5m
receivers:
- name: default
# your usual receivers
- name: hyperping-heartbeat
webhook_configs:
- url_file: /etc/alertmanager/secrets/hyperping-heartbeat/url
send_resolved: falseWhat each setting does here, from the Alertmanager configuration reference:
repeat_interval: 5mis the heartbeat rhythm. The docs say it "should be a multiple of the group_interval. If it's not, the repeat_interval is rounded up to the next multiple of the group_interval."group_interval: 1mis how often Alertmanager checks whether a repeat is due. A repeat fires on the first tick afterrepeat_intervalhas passed, so it can land one tick late.send_resolved: falsestops the "resolved" notification that would count as a ping (more on that below).
I measured the first point with repeat_interval: 1m and group_interval: 20s: pings arrived 80, 60, 80, 80 and 80 seconds apart. With 5m and 1m, expect 5 to 6 minutes, which the 5 minute period plus 5 minute grace covers.
Hyperping limits each healthcheck to 10 pings per minute. A single Watchdog group sends one ping per repeat, so you are far from it, even with several Alertmanager replicas.
3. Keep the ping URL in a file
The URL is the only credential of the healthcheck. url_file (Alertmanager 0.26 and later) reads it from a file, so it stays out of the config you commit:
echo -n 'https://hc.hyperping.io/tok_your_watchdog_token' > /etc/alertmanager/secrets/hyperping-heartbeat/urlWith kube-prometheus-stack, create a Secret and let the operator mount it. Secrets listed in alertmanagerSpec.secrets are mounted into /etc/alertmanager/secrets/<secret-name> in the Alertmanager container (Prometheus Operator API):
kubectl create secret generic hyperping-heartbeat -n monitoring \
--from-literal=url='https://hc.hyperping.io/tok_your_watchdog_token'# values.yaml for kube-prometheus-stack
alertmanager:
alertmanagerSpec:
secrets:
- hyperping-heartbeat
config:
route:
group_by: ['namespace']
group_wait: 30s
group_interval: 5m
repeat_interval: 12h
receiver: 'null'
routes:
- receiver: hyperping-heartbeat
matchers:
- alertname = "Watchdog"
group_wait: 0s
group_interval: 1m
repeat_interval: 5m
receivers:
- name: 'null'
- name: hyperping-heartbeat
webhook_configs:
- url_file: /etc/alertmanager/secrets/hyperping-heartbeat/url
send_resolved: falseHelm replaces lists instead of merging them, so this routes list replaces the chart's default one, including its Watchdog route to null. Add your other routes and receivers to the same lists.
I put this route in the main config rather than in an AlertmanagerConfig resource. By default, the operator adds a namespace matcher to every route that comes from an AlertmanagerConfig ("only process alerts that have a namespace label equal to the namespace of the object", per the operator's API), and vector(1) produces an alert without a namespace label, so such a route would not match it.
4. Check the route and the first pings
Before reloading Alertmanager, check the file and the route:
$ amtool check-config alertmanager.yml
Checking 'alertmanager.yml' SUCCESS
$ amtool config routes test --config.file=alertmanager.yml --verify.receivers=hyperping-heartbeat alertname=Watchdog severity=none
hyperping-heartbeatAfter the reload, open the healthcheck: the last pings should show POST requests with the user agent Alertmanager/0.34.1 (your version), and the healthcheck goes up.
Alertmanager's request is a POST with Content-Type: application/json and a JSON body ("version": "4", "status": "firing", the alerts). Hyperping's ping endpoint accepts HEAD, GET and POST, ignores the body, and answers 200, which counts as a successful notification. If the URL is wrong, it answers 401 and Alertmanager logs:
level=ERROR msg="Notify for alerts failed" ... err="hyperping-heartbeat/webhook[0]: notify retry canceled due to unrecoverable error after 1 attempts: unexpected status code 401: Token not valid"The webhook only retries 5xx responses and network errors; a 4xx is final. Since the failed notification is never recorded as sent, Alertmanager tries again at the next group_interval tick, so a fixed URL starts pinging within a minute.
5. Send the alerts somewhere that does not depend on Alertmanager
When the Watchdog pings stop, the alert comes from Hyperping, not from your cluster, so it arrives even if the cluster is gone. Alerts go to every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, so to wake up whoever is on call, connect PagerDuty or Opsgenie to the project. Avoid relying only on a channel that your Alertmanager also uses through the same integration: if that integration is what broke, you want a second path.
What happens when Prometheus dies
Prometheus does not tell Alertmanager that it stopped. Each alert it sends carries an end time, and in the rule evaluation code that end time is four times the larger of the evaluation interval and the resend delay ("Allow for two Eval or Alertmanager send failures"). With the default resend delay of 1 minute, Alertmanager keeps the Watchdog firing for up to 4 minutes after the last evaluation.
Here is what I recorded after killing Prometheus at 22:13:42, with two webhooks on the same route, one with send_resolved: false and one with the default:
| Time | send_resolved: false |
default (true) |
|---|---|---|
| 22:13:25 | firing | firing |
| 22:13:42 | Prometheus killed | Prometheus killed |
| 22:14:45 | firing | firing |
| 22:16:05 | firing | firing |
| 22:17:05 | nothing | resolved |
| after | nothing | nothing |
With the default, the last ping is the "resolved" notification, one cycle later. With send_resolved: false, the pings simply stop. Either way, count on the end time: with a 5 minute period and 5 minute grace, the alert comes roughly 10 to 15 minutes after Prometheus stops, and 10 minutes after the last ping when Alertmanager itself stops. For a faster alert, use repeat_interval: 2m, group_interval: 1m, and a 3 minute period with 3 minutes of grace.
Two more things stop the pings and are worth knowing:
- A silence that matches Watchdog. I added
amtool silence add alertname=Watchdog, and the pings stopped until I expired it. A broad maintenance silence (one matching every alert) will set off the dead man's switch, which is correct but surprising. - Inhibition rules. An
inhibit_rulesentry whose target matches Watchdog mutes it too.amtool alert query --inhibited alertname=Watchdoglists it when that happens.
In a highly available Alertmanager cluster, replicas share a notification log and only one of them sends each notification; the others wait their turn ("wait_time = peer_position × peer_timeout", in the HA docs). During a network partition, you may get duplicate pings, which is harmless.
Test it once by hand
Silence the Watchdog for 20 minutes (amtool silence add alertname=Watchdog --duration=20m --comment="dead man's switch test") and wait for the Hyperping alert after the period plus the grace period. Expire the silence: the next notification is a ping, and the incident resolves.
The same heartbeat covers the jobs your cluster runs on a schedule: see monitoring Kubernetes CronJobs, and database backups for the job you least want to fail silently. Node.js and Rails apps have their own guides: Node.js cron jobs and Rails scheduled jobs. For alert routing on the uptime side, my notes on alert management cover what I page on and what I don't.
FAQ
What is the Watchdog alert in Prometheus? ▼
Watchdog is an alert that is always firing, shipped by kube-prometheus and kube-prometheus-stack with the expression `vector(1)`. Its job is to prove that the whole alerting pipeline works: Prometheus evaluates rules, sends alerts to Alertmanager, and Alertmanager delivers notifications. You route it to an external dead man's switch, which alerts you when the notifications stop.
Why does kube-prometheus-stack route Watchdog to a null receiver? ▼
Because the chart cannot know where your dead man's switch lives. The default `alertmanager.config` has a route that sends `alertname = "Watchdog"` to the `null` receiver, so it fires but goes nowhere. To use it, replace that route with one pointing to a webhook receiver that calls your heartbeat URL.
Does a Hyperping healthcheck accept Alertmanager's webhook? ▼
Yes. Alertmanager's webhook sends a POST with a JSON body and the user agent `Alertmanager/<version>`. Hyperping healthcheck URLs accept HEAD, GET and POST, ignore the body, and answer 200, which Alertmanager counts as a successful notification. An invalid URL answers 401, which Alertmanager logs as an unrecoverable error.
What repeat_interval should the Watchdog route use? ▼
Short enough to detect an outage quickly, and a multiple of `group_interval`. I use `group_interval: 1m` and `repeat_interval: 5m`, with a healthcheck in simple mode every 5 minutes and a 5 minute grace period. In practice a repeat lands on a `group_interval` tick, so pings arrive every 5 to 6 minutes.
Should send_resolved be false on the Watchdog webhook? ▼
Yes. `send_resolved` defaults to true for webhooks. When Prometheus dies, Watchdog resolves in Alertmanager a few minutes later and Alertmanager sends a resolved notification, which a heartbeat URL counts as one more ping. In my test that delayed the alert by one more cycle. With `send_resolved: false`, the pings simply stop.
How long after Prometheus goes down does the dead man's switch alert? ▼
Prometheus sends alerts with an end time of four times the larger of the evaluation interval and the resend delay (1 minute by default), so Alertmanager keeps the Watchdog firing for about 3 to 4 minutes after the last evaluation. Add the healthcheck's period and grace period: with 5 and 5 minutes, expect the alert roughly 10 to 15 minutes after Prometheus stops.




