To monitor a scheduled AI agent, give it a heartbeat URL and call that URL at the end of the run, only after the agent's output passed a check you wrote. When a run is rate limited, loops until its timeout, runs out of budget, or "succeeds" with an empty answer, the call never happens and you get an alert. A process that exited with code 0 proves very little for an agent.
I'm Léo, I build Hyperping. A nightly eval, a triage bot or a weekly report writer is a cron job with a model loop inside, and it needs the same heartbeat as any other job, with one more check. This guide covers how these runs fail without anyone noticing, a pattern I tested for pinging only on a validated output, examples for four schedulers, and how an agent can check on the others over MCP.
Key takeaways
- Scheduled agents inherit every cron failure (the run never starts, the host is down) and add their own: 429 and 529 errors, loops cut by a turn or time limit, budget caps, refusals on HTTP 200, truncated output, revoked keys and retired models.
- Ping the heartbeat after you checked the result, not after the process exited. Check the result type (
subtypein Claude Code and the Agent SDK), a minimum size, and the fields only a real run fills. - Call
/startbefore the run so each run's duration is recorded, and size the grace period to the longest normal run plus your scheduler's delay: Claude Managed Agents adds up to 9 minutes of jitter, GitHub Actions can start an hour late. - Use one healthcheck per agent, in cron mode, with the same expression and timezone as its scheduler.
- Over the Hyperping MCP server,
list_healthcheckstells an agent which jobs are down, new or late, and when each next ping is due. That is enough for a morning brief or for an agent that refuses to run when its upstream job did not.
How scheduled AI agents fail silently
A scheduled agent is a cron job whose work is a model loop. It fails like any cron job, and it has a second set of failures that end with a clean exit and a useless result. I grouped them by what you see afterwards.
| Failure | What you see afterwards | Caught by a heartbeat? |
|---|---|---|
| Scheduler never starts the run | Nothing at all | Yes |
| Rate limit or overload past the retries | A non-zero exit, or nothing | Yes |
| Loop cut by max turns, budget or timeout | An error result, a killed process | Yes |
| Refusal, truncated or empty output | A clean exit and a bad result | Only if you check the output |
| Revoked key, no credits, retired model | The same error every night | Yes |
| Waiting for a tool approval nobody gives | A run that never ends | Yes, after the grace period |
The run never starts
This part is not specific to AI. GitHub Actions documents that scheduled runs can be delayed and that "some queued jobs may be dropped", and disables schedules in public repositories after 60 days without activity: I measured both in my guide to GitHub Actions scheduled workflows. Vercel calls cron delivery "best effort" and never retries a failed run (Vercel cron guide).
Claude Managed Agents scheduled deployments have their own cases. A session creation that is rate limited is recorded as a failed run with session_rate_limited_error and is not retried: "the schedule attempts again at the next scheduled occurrence". An archived environment or vault records a failed run and pauses the deployment, and an archived agent archives the deployment. After that, no session is created, so no agent can report anything.
The provider says no
The Claude API errors a nightly job runs into are 429 rate_limit_error, 529 overloaded_error and 500 api_error. The Anthropic SDKs retry those twice by default, which covers a blip, not a provider incident at 3 a.m. Other providers have the same classes of errors. When a provider is down for everyone, our Claude and OpenAI status checks show it, but your job still needs to tell you it did not run.
Then come the errors that repeat every night until someone looks: 401 authentication_error after a key was rotated or deleted, 402 billing_error for a billing problem such as prepaid credits running out, and 404 not_found_error when a hard-coded model ID was retired. The job fails the same way each run, quietly, until someone reads the log.
The loop runs out of turns, money or time
An agent decides how many steps it takes, so every runner puts a cap on it. In Claude Code and the Claude Agent SDK, a run that hits a cap ends with a result whose subtype is error_max_turns or error_max_budget_usd instead of success. A Managed Agents session that reaches its budget pauses with budget_reached. Your own timeout kills the process with exit code 124.
These caps are what you want: a runaway loop at night costs real money. The problem is that a capped run is easy to read as "it ran". The log has hundreds of lines of tool calls, and the last one is not the report.
The run "succeeds" with nothing in hand
This is the failure that exit codes miss. The stop reasons of the Messages API include max_tokens, a response cut off at the token limit, and refusal, a request the model declined, returned with HTTP 200. With structured output, the Agent SDK can end on error_max_structured_output_retries when the model could not produce valid JSON. And a model can write a polite paragraph saying the input folder was empty, with success as the result type, because an export upstream did not run.
None of this crashes. A script that does run_agent && curl $PING reports all of it as a success.
The run waits for an approval that never comes
In Managed Agents, MCP toolsets default to the always_ask permission policy: "The session pauses and waits for your approval before executing." A scheduled run has no client attached, so a run that reaches an MCP tool call without an always_allow policy, or without a webhook handler that answers it, sits there.
Why the agent's own logs are not enough
Every runner gives you a record of what happened:
- Claude Code with
--output-format jsonand the Agent SDK return a result withsubtype,is_error,num_turnsandtotal_cost_usd. - Managed Agents keeps a deployment run per trigger:
ant beta:deployment-runs list --deployment-id "$DEPLOYMENT_ID" --has-errorlists the failed ones. - GitHub Actions lists runs with
gh run list --workflow nightly-triage.yml --event schedule.
All of these describe runs that exist, and all of them need someone to look. A run that was never created, a deployment that paused itself, or a host that was down at 6:00 leaves no record to read. An external heartbeat flips the question: instead of looking for a failure, you wait for a success, and you hear about it when it does not come.
How to monitor scheduled AI agents with Hyperping
A Hyperping healthcheck is a secret URL, https://hc.hyperping.io/tok_..., that your job calls with HEAD, GET or POST. When the call does not arrive by the expected time plus a grace period, Hyperping opens an incident and alerts you, then resolves it on the next ping. Healthchecks are included on every plan, Free included, and each one counts as a monitor.
1. Create a healthcheck with the agent's schedule and timezone
In Hyperping, open Healthchecks, click Create healthcheck, name it after the agent, and pick Cron. Copy the expression from the scheduler that starts the agent and pick the same IANA timezone. For a support digest that cron starts at 6:00 Paris time on weekdays:
- Cron expression:
0 6 * * 1-5(every weekday, at 6:00 instead of midnight) - Timezone:
Europe/Paris
Hyperping then expects a ping at 04:00 UTC in October and at 05:00 UTC after the clocks change, like the cron that starts the job. For agents that run in a loop rather than on a clock, simple mode (every N minutes, hours or days) works too. The healthchecks docs cover both modes.
2. Ping /start when the run begins
Call the healthcheck URL followed by /start before the agent starts. It opens a run without moving the deadline, and the next success ping closes it and records the duration. After a week, the healthcheck's last pings show how long each run took: an agent whose runs grow from 6 to 14 minutes is telling you something before it hits its timeout.
3. Ping only after the output passes a check
This is the part that differs from a regular cron job. Here is a wrapper around Claude Code in headless mode for the support digest. It asks for a structured result with --json-schema, caps turns, money and time, and checks the result before it pings:
#!/usr/bin/env bash
# Nightly support digest, run by cron at 06:00 Paris time on weekdays.
set -uo pipefail
HC="${HYPERPING_DIGEST_URL:?set it to the healthcheck URL}"
OUT="$(mktemp)"
trap 'rm -f "$OUT"' EXIT
ping() {
curl -fsS -m 10 --retry 3 -o /dev/null "$1" || echo "ping failed: $1" >&2
}
ping "$HC/start"
timeout 20m claude -p "$(cat prompts/digest.md)" \
--output-format json \
--json-schema "$(cat schemas/digest.json)" \
--max-turns 40 \
--max-budget-usd 3 \
--allowedTools "Read" "Grep" "Glob" \
> "$OUT"
status=$?
if [ "$status" -ne 0 ]; then
echo "agent exited with $status (124 = killed by timeout)" >&2
exit "$status"
fi
# Ping only when the output is the real thing, not when the process exited.
if jq -e '
.subtype == "success"
and (.structured_output.tickets_read | type == "number" and . > 0)
and (.structured_output.urgent | type == "array")
and (.structured_output.summary | length >= 200)
' "$OUT" > /dev/null; then
jq '.structured_output' "$OUT" > "digest-$(date +%F).json"
ping "$HC"
else
echo "output failed validation:" >&2
jq -c '{subtype, is_error, structured_output}' "$OUT" >&2
exit 1
fischemas/digest.json asks for tickets_read (an integer), urgent (a list of ticket IDs with a reason) and summary (a string). The checks are the ones that hold on every good run of this agent: on a weekday, zero tickets read means the ticket export broke, and a summary under 200 characters is an apology, not a digest. Pick yours the same way. A count of items processed, a date field equal to today, a file over a minimum size, or a section count in a report are all cheap to check.
I tested the wrapper on macOS with a stand-in claude command and a local listener in place of hc.hyperping.io, without calling the API:
| Stand-in behavior | Exit code | Pings received |
|---|---|---|
success with 42 tickets and a 250 character summary |
0 | /start, then the success ping |
success with 0 tickets and "No data." |
1 | /start only |
error_max_turns, exit 1 |
1 | /start only |
429 rate_limit_error on stderr, exit 1 |
1 | /start only |
| Never returns (timeout set to 3 seconds for the test) | 124 | /start only |
The ping function never fails the job: curl gives up after 10 seconds per try and the error goes to stderr. On Linux, timeout comes with coreutils. On macOS, install it with brew install coreutils.
4. Size the grace period to the run and the scheduler
In cron mode, Hyperping marks the healthcheck down when the success ping has not arrived by the scheduled time plus the grace period, and the success ping leaves at the end of the run. So the grace period has to cover:
| Scheduler | Delay it adds | Grace period I would start with |
|---|---|---|
cron or systemd on your server, 20 minute timeout |
None | 25 minutes |
GitHub Actions, daily, timeout-minutes: 30 |
Usually 3 to 15 minutes, once 1 h 39 in October 2026 | 2 hours |
| Managed Agents deployment, run of up to 25 minutes | Jitter up to 15% of the interval, capped at 9 minutes | 45 minutes |
| Vercel Cron on Hobby | Anywhere in the scheduled hour | 75 minutes for a 300 second function |
After a week of runs, tighten it using the durations in the ping history. The default is 10 minutes, the minimum 1 minute.
5. Route the alerts
A missed run alerts every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, those are for uptime monitors. If a missed eval run before a release has to wake someone up, send healthcheck alerts to PagerDuty or Opsgenie and let their rotation decide who gets paged. The incident resolves on its own at the next successful ping.
Examples for four other ways to schedule an agent
The pattern stays the same everywhere: /start first, the agent, a check on its result, then the success ping. Only where the check runs changes.
A Claude Agent SDK script started by cron
With the Agent SDK (pip install claude-agent-sdk), the final ResultMessage carries the subtype and the structured_output:
import asyncio
import json
import os
import sys
import urllib.request
from datetime import date
from claude_agent_sdk import ClaudeAgentOptions, ResultMessage, query
HC = os.environ["HYPERPING_DIGEST_URL"] # https://hc.hyperping.io/tok_...
SCHEMA = {
"type": "object",
"properties": {
"tickets_read": {"type": "integer"},
"urgent": {
"type": "array",
"items": {
"type": "object",
"properties": {"id": {"type": "string"}, "reason": {"type": "string"}},
"required": ["id", "reason"],
},
},
"summary": {"type": "string"},
},
"required": ["tickets_read", "urgent", "summary"],
}
def ping(url: str) -> None:
# A monitoring call must never be the reason the job fails.
try:
urllib.request.urlopen(url, timeout=10).read()
except Exception as error:
print(f"ping failed: {url}: {error}", file=sys.stderr)
def looks_real(output) -> bool:
return (
isinstance(output, dict)
and isinstance(output.get("tickets_read"), int)
and output["tickets_read"] > 0
and isinstance(output.get("urgent"), list)
and len(output.get("summary", "")) >= 200
)
async def main() -> int:
ping(f"{HC}/start")
options = ClaudeAgentOptions(
model="claude-opus-5-5",
cwd="/srv/support-export",
allowed_tools=["Read", "Grep", "Glob"],
max_turns=40,
max_budget_usd=3.0,
output_format={"type": "json_schema", "schema": SCHEMA},
)
result = None
try:
async for message in query(
prompt="Read yesterday's tickets in ./tickets and write the morning digest.",
options=options,
):
if isinstance(message, ResultMessage):
result = message
except Exception as error: # CLI crash, auth error, network error
print(f"agent failed: {error}", file=sys.stderr)
return 1
if result is None or result.subtype != "success":
# error_max_turns, error_max_budget_usd, error_during_execution...
print(f"run ended with {result and result.subtype}", file=sys.stderr)
return 1
if not looks_real(result.structured_output):
print(f"output failed validation: {result.structured_output!r}", file=sys.stderr)
return 1
with open(f"digest-{date.today()}.json", "w") as f:
json.dump(result.structured_output, f, indent=2)
ping(HC)
return 0
if __name__ == "__main__":
sys.exit(asyncio.run(main()))I checked it against claude-agent-sdk 0.2.165 on Python 3.14, with query replaced by a stub: a success result with a real digest sent /start then the success ping, and a result with zero tickets, an error_max_budget_usd result, an exception, or no result at all sent /start only and exited 1. Start it from cron with 0 6 * * 1-5, and CRON_TZ=Europe/Paris on cronie (otherwise the server's timezone applies), the same values as the healthcheck.
Claude Code in a scheduled GitHub Actions workflow
anthropics/claude-code-action@v1 runs Claude Code in a workflow. With --json-schema in claude_args, the step exposes the result as structured_output, and a later step can check it before pinging:
name: Nightly issue triage
on:
schedule:
- cron: "30 5 * * 1-5"
timezone: "Europe/Paris"
workflow_dispatch:
jobs:
triage:
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: read
issues: write
id-token: write
env:
HC: ${{ secrets.HYPERPING_TRIAGE_URL }}
steps:
- name: Ping start
run: curl -fsS -m 10 --retry 3 -o /dev/null "$HC/start" || true
- uses: actions/checkout@v6
- id: triage
uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
prompt: |
REPO: ${{ github.repository }}
Label every issue opened in the last 24 hours as bug, question or feature,
and list the ones that need a human today.
claude_args: |
--max-turns 40
--allowedTools "Bash(gh issue:*)"
--json-schema '{"type":"object","properties":{"issues_seen":{"type":"integer"},"labeled":{"type":"integer"},"needs_human":{"type":"array","items":{"type":"integer"}}},"required":["issues_seen","labeled","needs_human"]}'
- name: Check the output, then ping
env:
OUTPUT: ${{ steps.triage.outputs.structured_output }}
run: |
echo "$OUTPUT" | jq -e '
(.issues_seen | type == "number")
and (.labeled == .issues_seen)
and (.needs_human | type == "array")
' > /dev/null
curl -fsS -m 10 --retry 3 -o /dev/null "$HC"The last step only runs when every step before it succeeded, and jq -e fails it when an issue was left unlabeled or the output is missing. I ran that check on a complete output (exit 0), on one with 3 of 5 issues labeled (exit 1) and on an empty one (exit 4). The timezone key next to cron exists since March 2026: set the healthcheck to 30 5 * * 1-5 in Europe/Paris too. I did not run this workflow on GitHub; the YAML and the schema were parsed locally.
A Claude Managed Agents scheduled deployment
A scheduled deployment runs on Anthropic's side, so the pings have to come from the session's sandbox. Two settings make that work. First, an environment whose limited networking allows the healthcheck host:
# environment.yaml
name: scheduled-agents
config:
type: cloud
networking:
type: limited
allowed_hosts:
- hc.hyperping.ioThen a deployment whose instructions end on one command that checks the result and pings only if the check passes:
---
name: Weekly dependency report
agent: agent_01...
environment_id: env_01...
schedule:
type: cron
expression: "0 7 * * 1"
timezone: Europe/Paris
---
Start by running: curl -fsS -m 10 --retry 3 https://hc.hyperping.io/tok_.../start
Review last week's dependency updates and write report.md, with one ## section per service.
To finish, run exactly this command and report its exit code:
test "$(wc -c < report.md)" -ge 2000 && test "$(grep -c '^## ' report.md)" -ge 3 && curl -fsS -m 10 --retry 3 https://hc.hyperping.io/tok_...Sync both with ant apply. The && chain is what keeps the check out of the model's hands: the ping only leaves when the file exists, is at least 2,000 bytes and has 3 sections. I ran that chain locally on a complete report, a one-line report and a missing file: only the first passed. Every case where no session starts (rate limit, paused or archived deployment) also means no ping, which is the point.
Two settings to check. If the agent uses MCP tools, give them an always_allow permission policy, or the run waits for an approval no one is there to give. And set a budget on the deployment: it is copied onto each session, so one bad run cannot spend a month's money. The ping token sits in the deployment's instructions and in the session transcript, so treat it as a secret: anyone who has it can send pings, though it gives no read access to anything.
An agent behind Vercel Cron
If a Vercel Cron route calls the Claude API, the Vercel cron guide covers the route, CRON_SECRET and the ping helper. Two things change for an agent. Ping only after the response passed your check (with client.messages.parse() and structured output, parsed_output must be non-null and stop_reason must be end_turn). And watch maxDuration: on Hobby a function stops at 300 seconds, which a multi-step agent loop can reach, and Vercel does not retry it.
Let your agents check on each other over MCP
Alerts tell you when a run is missing. The Hyperping MCP server lets an agent answer the question you ask over coffee: did last night's jobs run? It needs an API key (Essentials, Pro and Business plans), and a read-only key is enough for everything below. Setup for Claude Code, Cursor, Codex and others is in the agent setup docs.
Two tools cover healthchecks:
list_healthchecksreturns every healthcheck with itsstate(down,up, ornewfor never pinged),last_ping_at,schedule,grace, and for the ones that are down,down_sinceand the ongoing incident. Up ones also getnext_ping_due_at,down_if_no_ping_by, andlate: truewhile they are inside their grace period. These two times are not shown in the dashboard.get_healthcheckadds the last 20 pings (kept 15 days) with each run's duration when/startwas used, uptime and downtime over 30 days, the hosts that ping it, and the last 10 incidents.
A morning brief of late jobs
At 7:15 Paris time, after the digest and the triage, a third scheduled run asks Claude Code to read the healthchecks. --mcp-config accepts a JSON string, so the key stays in an environment variable:
MCP_CONFIG=$(jq -n --arg key "$HYPERPING_API_KEY" '{
mcpServers: { hyperping: {
type: "http",
url: "https://api.hyperping.io/v1/mcp",
headers: { Authorization: ("Bearer " + $key) }
} }
}')
claude -p "List my Hyperping healthchecks that are down, new or late. For each down one, \
give the last ping and the last 3 runs from get_healthcheck. Ten lines at most." \
--mcp-config "$MCP_CONFIG" \
--strict-mcp-config \
--allowedTools "mcp__hyperping__list_healthchecks" "mcp__hyperping__get_healthcheck"On the morning the digest failed, list_healthchecks returns something like this (names and IDs shortened):
{
"counts": { "total": 4, "down": 1, "up": 3, "new": 0 },
"returned": 4,
"truncated": false,
"healthchecks": [
{
"uuid": "hc_...",
"name": "Nightly support digest",
"state": "down",
"last_ping_at": "2026-10-07T04:03:12Z",
"down_since": "2026-10-08T04:25:00Z",
"schedule": "cron '0 6 * * 1-5' (Europe/Paris)",
"grace": "25 minutes",
"ongoing_incident": "hco_..."
},
{
"uuid": "hc_...",
"name": "Eval suite",
"state": "up",
"last_ping_at": "2026-10-07T03:48:40Z",
"next_ping_due_at": "2026-10-08T03:00:00Z",
"down_if_no_ping_by": "2026-10-08T06:00:00Z",
"late": true,
"schedule": "cron '0 5 * * *' (Europe/Paris)",
"grace": "3 hours"
}
]
}The brief can then say: the support digest has been down since 6:25, its last good run was yesterday at 6:03, and the eval suite is 2 hours 15 minutes late but has until 8:00 before it alerts. Send the text wherever your team reads in the morning, and give the brief its own healthcheck: it is a scheduled agent too.
The --allowedTools list lets the two read tools run without a prompt, and --strict-mcp-config keeps other configured servers out of this run. With a read-only key, write tools are refused by the server anyway.
An agent that checks its upstream job first
The same tools stop an agent from working on stale input. The digest reads an export that another job writes at 5:30. Add one line at the top of its prompt:
Before anything else, call get_healthcheck for "Support export". If its state is not up, or its last_ping_at is earlier than today 03:30 UTC, stop and say the export did not run.
The digest then fails its own output check, its healthcheck goes down, and the alert names the real cause instead of an empty summary. The export has its own healthcheck, so both alerts arrive, in order.
The MCP server allows 60 tool calls a minute and 600 an hour per project. A brief that reads a few dozen healthchecks once a day is far from that.
Where to start
Pick the scheduled agent whose silence would cost you the most, often the eval run before a release or the job that feeds a customer-facing report. Give it a healthcheck in cron mode with its own schedule and timezone, wrap the run in /start and a checked success ping, and let a week of durations tell you the right grace period. Then add the morning brief.
For the plain cron jobs around your agents, the same approach is in the guides for Node.js cron jobs, GitHub Actions scheduled workflows and Vercel Cron, and Hyperping for AI companies covers the rest of the stack: inference endpoints, model providers and status pages.
FAQ
How do I know if my scheduled AI agent ran? ▼
Give each scheduled agent its own heartbeat URL and call it at the end of the run, only after the output passed a check. A heartbeat monitor in cron mode, with the same expression and timezone as the agent's scheduler, alerts you when the call does not arrive by the scheduled time plus a grace period. Your scheduler's logs only show runs that started, not the ones that never did.
Why is an exit code of 0 not enough for an AI agent? ▼
Because an agent can finish cleanly with nothing useful in hand. A request can end on a refusal with HTTP 200, an answer can stop at max_tokens, a structured output can fail its schema retries, and a model can write a short apology instead of the report. Check the result itself, for example its subtype, a minimum size and the fields only a real run fills, before you report success.
How long should the grace period be for an AI agent heartbeat? ▼
The longest normal run, plus the delay your scheduler adds, plus a margin. A script with a 20 minute timeout started by cron needs about 25 minutes. A GitHub Actions schedule can start more than an hour late, so give a daily workflow one to two hours. Claude Managed Agents adds up to 9 minutes of jitter to each scheduled run.
Can an AI agent check whether my other cron jobs ran? ▼
Yes. Connect it to the Hyperping MCP server with a read-only API key. The list_healthchecks tool returns each healthcheck's state, last ping, the time the next ping is due, the time it will be marked down, and a late flag while it is in its grace period. get_healthcheck adds the last 20 pings with run durations and recent incidents.
Can Hyperping page my on-call engineer when a scheduled agent misses a run? ▼
Healthchecks alert every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. They do not use Hyperping's escalation policies or on-call schedules, so to wake someone up, send the alert to PagerDuty or Opsgenie and let their rotation decide who gets paged.




