To monitor a scheduled AI agent, give it a heartbeat URL and call that URL at the end of the run, only after the agent's output passed a check you wrote. When a run is rate limited, loops until its timeout, runs out of budget, or "succeeds" with an empty answer, the call never happens and you get an alert. A process that exited with code 0 proves very little for an agent.

I'm Léo, I build Hyperping. A nightly eval, a triage bot or a weekly report writer is a cron job with a model loop inside, and it needs the same heartbeat as any other job, with one more check. This guide covers how these runs fail without anyone noticing, a pattern I tested for pinging only on a validated output, examples for four schedulers, and how an agent can check on the others over MCP.

Key takeaways

  • Scheduled agents inherit every cron failure (the run never starts, the host is down) and add their own: 429 and 529 errors, loops cut by a turn or time limit, budget caps, refusals on HTTP 200, truncated output, revoked keys and retired models.
  • Ping the heartbeat after you checked the result, not after the process exited. Check the result type (subtype in Claude Code and the Agent SDK), a minimum size, and the fields only a real run fills.
  • Call /start before the run so each run's duration is recorded, and size the grace period to the longest normal run plus your scheduler's delay: Claude Managed Agents adds up to 9 minutes of jitter, GitHub Actions can start an hour late.
  • Use one healthcheck per agent, in cron mode, with the same expression and timezone as its scheduler.
  • Over the Hyperping MCP server, list_healthchecks tells an agent which jobs are down, new or late, and when each next ping is due. That is enough for a morning brief or for an agent that refuses to run when its upstream job did not.

How scheduled AI agents fail silently

A scheduled agent is a cron job whose work is a model loop. It fails like any cron job, and it has a second set of failures that end with a clean exit and a useless result. I grouped them by what you see afterwards.

Failure What you see afterwards Caught by a heartbeat?
Scheduler never starts the run Nothing at all Yes
Rate limit or overload past the retries A non-zero exit, or nothing Yes
Loop cut by max turns, budget or timeout An error result, a killed process Yes
Refusal, truncated or empty output A clean exit and a bad result Only if you check the output
Revoked key, no credits, retired model The same error every night Yes
Waiting for a tool approval nobody gives A run that never ends Yes, after the grace period

The run never starts

This part is not specific to AI. GitHub Actions documents that scheduled runs can be delayed and that "some queued jobs may be dropped", and disables schedules in public repositories after 60 days without activity: I measured both in my guide to GitHub Actions scheduled workflows. Vercel calls cron delivery "best effort" and never retries a failed run (Vercel cron guide).

Claude Managed Agents scheduled deployments have their own cases. A session creation that is rate limited is recorded as a failed run with session_rate_limited_error and is not retried: "the schedule attempts again at the next scheduled occurrence". An archived environment or vault records a failed run and pauses the deployment, and an archived agent archives the deployment. After that, no session is created, so no agent can report anything.

The provider says no

The Claude API errors a nightly job runs into are 429 rate_limit_error, 529 overloaded_error and 500 api_error. The Anthropic SDKs retry those twice by default, which covers a blip, not a provider incident at 3 a.m. Other providers have the same classes of errors. When a provider is down for everyone, our Claude and OpenAI status checks show it, but your job still needs to tell you it did not run.

Then come the errors that repeat every night until someone looks: 401 authentication_error after a key was rotated or deleted, 402 billing_error for a billing problem such as prepaid credits running out, and 404 not_found_error when a hard-coded model ID was retired. The job fails the same way each run, quietly, until someone reads the log.

The loop runs out of turns, money or time

An agent decides how many steps it takes, so every runner puts a cap on it. In Claude Code and the Claude Agent SDK, a run that hits a cap ends with a result whose subtype is error_max_turns or error_max_budget_usd instead of success. A Managed Agents session that reaches its budget pauses with budget_reached. Your own timeout kills the process with exit code 124.

These caps are what you want: a runaway loop at night costs real money. The problem is that a capped run is easy to read as "it ran". The log has hundreds of lines of tool calls, and the last one is not the report.

The run "succeeds" with nothing in hand

This is the failure that exit codes miss. The stop reasons of the Messages API include max_tokens, a response cut off at the token limit, and refusal, a request the model declined, returned with HTTP 200. With structured output, the Agent SDK can end on error_max_structured_output_retries when the model could not produce valid JSON. And a model can write a polite paragraph saying the input folder was empty, with success as the result type, because an export upstream did not run.

None of this crashes. A script that does run_agent && curl $PING reports all of it as a success.

The run waits for an approval that never comes

In Managed Agents, MCP toolsets default to the always_ask permission policy: "The session pauses and waits for your approval before executing." A scheduled run has no client attached, so a run that reaches an MCP tool call without an always_allow policy, or without a webhook handler that answers it, sits there.

Why the agent's own logs are not enough

Every runner gives you a record of what happened:

  • Claude Code with --output-format json and the Agent SDK return a result with subtype, is_error, num_turns and total_cost_usd.
  • Managed Agents keeps a deployment run per trigger: ant beta:deployment-runs list --deployment-id "$DEPLOYMENT_ID" --has-error lists the failed ones.
  • GitHub Actions lists runs with gh run list --workflow nightly-triage.yml --event schedule.

All of these describe runs that exist, and all of them need someone to look. A run that was never created, a deployment that paused itself, or a host that was down at 6:00 leaves no record to read. An external heartbeat flips the question: instead of looking for a failure, you wait for a success, and you hear about it when it does not come.

How to monitor scheduled AI agents with Hyperping

A Hyperping healthcheck is a secret URL, https://hc.hyperping.io/tok_..., that your job calls with HEAD, GET or POST. When the call does not arrive by the expected time plus a grace period, Hyperping opens an incident and alerts you, then resolves it on the next ping. Healthchecks are included on every plan, Free included, and each one counts as a monitor.

1. Create a healthcheck with the agent's schedule and timezone

In Hyperping, open Healthchecks, click Create healthcheck, name it after the agent, and pick Cron. Copy the expression from the scheduler that starts the agent and pick the same IANA timezone. For a support digest that cron starts at 6:00 Paris time on weekdays:

  • Cron expression: 0 6 * * 1-5 (every weekday, at 6:00 instead of midnight)
  • Timezone: Europe/Paris

Hyperping then expects a ping at 04:00 UTC in October and at 05:00 UTC after the clocks change, like the cron that starts the job. For agents that run in a loop rather than on a clock, simple mode (every N minutes, hours or days) works too. The healthchecks docs cover both modes.

2. Ping /start when the run begins

Call the healthcheck URL followed by /start before the agent starts. It opens a run without moving the deadline, and the next success ping closes it and records the duration. After a week, the healthcheck's last pings show how long each run took: an agent whose runs grow from 6 to 14 minutes is telling you something before it hits its timeout.

3. Ping only after the output passes a check

This is the part that differs from a regular cron job. Here is a wrapper around Claude Code in headless mode for the support digest. It asks for a structured result with --json-schema, caps turns, money and time, and checks the result before it pings:

#!/usr/bin/env bash
# Nightly support digest, run by cron at 06:00 Paris time on weekdays.
set -uo pipefail

HC="${HYPERPING_DIGEST_URL:?set it to the healthcheck URL}"
OUT="$(mktemp)"
trap 'rm -f "$OUT"' EXIT

ping() {
  curl -fsS -m 10 --retry 3 -o /dev/null "$1" || echo "ping failed: $1" >&2
}

ping "$HC/start"

timeout 20m claude -p "$(cat prompts/digest.md)" \
  --output-format json \
  --json-schema "$(cat schemas/digest.json)" \
  --max-turns 40 \
  --max-budget-usd 3 \
  --allowedTools "Read" "Grep" "Glob" \
  > "$OUT"
status=$?

if [ "$status" -ne 0 ]; then
  echo "agent exited with $status (124 = killed by timeout)" >&2
  exit "$status"
fi

# Ping only when the output is the real thing, not when the process exited.
if jq -e '
  .subtype == "success"
  and (.structured_output.tickets_read | type == "number" and . > 0)
  and (.structured_output.urgent | type == "array")
  and (.structured_output.summary | length >= 200)
' "$OUT" > /dev/null; then
  jq '.structured_output' "$OUT" > "digest-$(date +%F).json"
  ping "$HC"
else
  echo "output failed validation:" >&2
  jq -c '{subtype, is_error, structured_output}' "$OUT" >&2
  exit 1
fi

schemas/digest.json asks for tickets_read (an integer), urgent (a list of ticket IDs with a reason) and summary (a string). The checks are the ones that hold on every good run of this agent: on a weekday, zero tickets read means the ticket export broke, and a summary under 200 characters is an apology, not a digest. Pick yours the same way. A count of items processed, a date field equal to today, a file over a minimum size, or a section count in a report are all cheap to check.

I tested the wrapper on macOS with a stand-in claude command and a local listener in place of hc.hyperping.io, without calling the API:

Stand-in behavior Exit code Pings received
success with 42 tickets and a 250 character summary 0 /start, then the success ping
success with 0 tickets and "No data." 1 /start only
error_max_turns, exit 1 1 /start only
429 rate_limit_error on stderr, exit 1 1 /start only
Never returns (timeout set to 3 seconds for the test) 124 /start only

The ping function never fails the job: curl gives up after 10 seconds per try and the error goes to stderr. On Linux, timeout comes with coreutils. On macOS, install it with brew install coreutils.

4. Size the grace period to the run and the scheduler

In cron mode, Hyperping marks the healthcheck down when the success ping has not arrived by the scheduled time plus the grace period, and the success ping leaves at the end of the run. So the grace period has to cover:

Scheduler Delay it adds Grace period I would start with
cron or systemd on your server, 20 minute timeout None 25 minutes
GitHub Actions, daily, timeout-minutes: 30 Usually 3 to 15 minutes, once 1 h 39 in October 2026 2 hours
Managed Agents deployment, run of up to 25 minutes Jitter up to 15% of the interval, capped at 9 minutes 45 minutes
Vercel Cron on Hobby Anywhere in the scheduled hour 75 minutes for a 300 second function

After a week of runs, tighten it using the durations in the ping history. The default is 10 minutes, the minimum 1 minute.

5. Route the alerts

A missed run alerts every channel connected to the project at once: email and SMS to every member, Slack, Discord, Telegram, PagerDuty and Opsgenie. Healthchecks do not use escalation policies or on-call schedules, those are for uptime monitors. If a missed eval run before a release has to wake someone up, send healthcheck alerts to PagerDuty or Opsgenie and let their rotation decide who gets paged. The incident resolves on its own at the next successful ping.

Examples for four other ways to schedule an agent

The pattern stays the same everywhere: /start first, the agent, a check on its result, then the success ping. Only where the check runs changes.

A Claude Agent SDK script started by cron

With the Agent SDK (pip install claude-agent-sdk), the final ResultMessage carries the subtype and the structured_output:

import asyncio
import json
import os
import sys
import urllib.request
from datetime import date

from claude_agent_sdk import ClaudeAgentOptions, ResultMessage, query

HC = os.environ["HYPERPING_DIGEST_URL"]  # https://hc.hyperping.io/tok_...

SCHEMA = {
    "type": "object",
    "properties": {
        "tickets_read": {"type": "integer"},
        "urgent": {
            "type": "array",
            "items": {
                "type": "object",
                "properties": {"id": {"type": "string"}, "reason": {"type": "string"}},
                "required": ["id", "reason"],
            },
        },
        "summary": {"type": "string"},
    },
    "required": ["tickets_read", "urgent", "summary"],
}


def ping(url: str) -> None:
    # A monitoring call must never be the reason the job fails.
    try:
        urllib.request.urlopen(url, timeout=10).read()
    except Exception as error:
        print(f"ping failed: {url}: {error}", file=sys.stderr)


def looks_real(output) -> bool:
    return (
        isinstance(output, dict)
        and isinstance(output.get("tickets_read"), int)
        and output["tickets_read"] > 0
        and isinstance(output.get("urgent"), list)
        and len(output.get("summary", "")) >= 200
    )


async def main() -> int:
    ping(f"{HC}/start")

    options = ClaudeAgentOptions(
        model="claude-opus-5-5",
        cwd="/srv/support-export",
        allowed_tools=["Read", "Grep", "Glob"],
        max_turns=40,
        max_budget_usd=3.0,
        output_format={"type": "json_schema", "schema": SCHEMA},
    )

    result = None
    try:
        async for message in query(
            prompt="Read yesterday's tickets in ./tickets and write the morning digest.",
            options=options,
        ):
            if isinstance(message, ResultMessage):
                result = message
    except Exception as error:  # CLI crash, auth error, network error
        print(f"agent failed: {error}", file=sys.stderr)
        return 1

    if result is None or result.subtype != "success":
        # error_max_turns, error_max_budget_usd, error_during_execution...
        print(f"run ended with {result and result.subtype}", file=sys.stderr)
        return 1

    if not looks_real(result.structured_output):
        print(f"output failed validation: {result.structured_output!r}", file=sys.stderr)
        return 1

    with open(f"digest-{date.today()}.json", "w") as f:
        json.dump(result.structured_output, f, indent=2)

    ping(HC)
    return 0


if __name__ == "__main__":
    sys.exit(asyncio.run(main()))

I checked it against claude-agent-sdk 0.2.165 on Python 3.14, with query replaced by a stub: a success result with a real digest sent /start then the success ping, and a result with zero tickets, an error_max_budget_usd result, an exception, or no result at all sent /start only and exited 1. Start it from cron with 0 6 * * 1-5, and CRON_TZ=Europe/Paris on cronie (otherwise the server's timezone applies), the same values as the healthcheck.

Claude Code in a scheduled GitHub Actions workflow

anthropics/claude-code-action@v1 runs Claude Code in a workflow. With --json-schema in claude_args, the step exposes the result as structured_output, and a later step can check it before pinging:

name: Nightly issue triage

on:
  schedule:
    - cron: "30 5 * * 1-5"
      timezone: "Europe/Paris"
  workflow_dispatch:

jobs:
  triage:
    runs-on: ubuntu-latest
    timeout-minutes: 30
    permissions:
      contents: read
      issues: write
      id-token: write
    env:
      HC: ${{ secrets.HYPERPING_TRIAGE_URL }}
    steps:
      - name: Ping start
        run: curl -fsS -m 10 --retry 3 -o /dev/null "$HC/start" || true

      - uses: actions/checkout@v6

      - id: triage
        uses: anthropics/claude-code-action@v1
        with:
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
          prompt: |
            REPO: ${{ github.repository }}
            Label every issue opened in the last 24 hours as bug, question or feature,
            and list the ones that need a human today.
          claude_args: |
            --max-turns 40
            --allowedTools "Bash(gh issue:*)"
            --json-schema '{"type":"object","properties":{"issues_seen":{"type":"integer"},"labeled":{"type":"integer"},"needs_human":{"type":"array","items":{"type":"integer"}}},"required":["issues_seen","labeled","needs_human"]}'

      - name: Check the output, then ping
        env:
          OUTPUT: ${{ steps.triage.outputs.structured_output }}
        run: |
          echo "$OUTPUT" | jq -e '
            (.issues_seen | type == "number")
            and (.labeled == .issues_seen)
            and (.needs_human | type == "array")
          ' > /dev/null
          curl -fsS -m 10 --retry 3 -o /dev/null "$HC"

The last step only runs when every step before it succeeded, and jq -e fails it when an issue was left unlabeled or the output is missing. I ran that check on a complete output (exit 0), on one with 3 of 5 issues labeled (exit 1) and on an empty one (exit 4). The timezone key next to cron exists since March 2026: set the healthcheck to 30 5 * * 1-5 in Europe/Paris too. I did not run this workflow on GitHub; the YAML and the schema were parsed locally.

A Claude Managed Agents scheduled deployment

A scheduled deployment runs on Anthropic's side, so the pings have to come from the session's sandbox. Two settings make that work. First, an environment whose limited networking allows the healthcheck host:

# environment.yaml
name: scheduled-agents
config:
  type: cloud
  networking:
    type: limited
    allowed_hosts:
      - hc.hyperping.io

Then a deployment whose instructions end on one command that checks the result and pings only if the check passes:

---
name: Weekly dependency report
agent: agent_01...
environment_id: env_01...
schedule:
  type: cron
  expression: "0 7 * * 1"
  timezone: Europe/Paris
---

Start by running: curl -fsS -m 10 --retry 3 https://hc.hyperping.io/tok_.../start

Review last week's dependency updates and write report.md, with one ## section per service.

To finish, run exactly this command and report its exit code:
test "$(wc -c < report.md)" -ge 2000 && test "$(grep -c '^## ' report.md)" -ge 3 && curl -fsS -m 10 --retry 3 https://hc.hyperping.io/tok_...

Sync both with ant apply. The && chain is what keeps the check out of the model's hands: the ping only leaves when the file exists, is at least 2,000 bytes and has 3 sections. I ran that chain locally on a complete report, a one-line report and a missing file: only the first passed. Every case where no session starts (rate limit, paused or archived deployment) also means no ping, which is the point.

Two settings to check. If the agent uses MCP tools, give them an always_allow permission policy, or the run waits for an approval no one is there to give. And set a budget on the deployment: it is copied onto each session, so one bad run cannot spend a month's money. The ping token sits in the deployment's instructions and in the session transcript, so treat it as a secret: anyone who has it can send pings, though it gives no read access to anything.

An agent behind Vercel Cron

If a Vercel Cron route calls the Claude API, the Vercel cron guide covers the route, CRON_SECRET and the ping helper. Two things change for an agent. Ping only after the response passed your check (with client.messages.parse() and structured output, parsed_output must be non-null and stop_reason must be end_turn). And watch maxDuration: on Hobby a function stops at 300 seconds, which a multi-step agent loop can reach, and Vercel does not retry it.

Let your agents check on each other over MCP

Alerts tell you when a run is missing. The Hyperping MCP server lets an agent answer the question you ask over coffee: did last night's jobs run? It needs an API key (Essentials, Pro and Business plans), and a read-only key is enough for everything below. Setup for Claude Code, Cursor, Codex and others is in the agent setup docs.

Two tools cover healthchecks:

  • list_healthchecks returns every healthcheck with its state (down, up, or new for never pinged), last_ping_at, schedule, grace, and for the ones that are down, down_since and the ongoing incident. Up ones also get next_ping_due_at, down_if_no_ping_by, and late: true while they are inside their grace period. These two times are not shown in the dashboard.
  • get_healthcheck adds the last 20 pings (kept 15 days) with each run's duration when /start was used, uptime and downtime over 30 days, the hosts that ping it, and the last 10 incidents.

A morning brief of late jobs

At 7:15 Paris time, after the digest and the triage, a third scheduled run asks Claude Code to read the healthchecks. --mcp-config accepts a JSON string, so the key stays in an environment variable:

MCP_CONFIG=$(jq -n --arg key "$HYPERPING_API_KEY" '{
  mcpServers: { hyperping: {
    type: "http",
    url: "https://api.hyperping.io/v1/mcp",
    headers: { Authorization: ("Bearer " + $key) }
  } }
}')

claude -p "List my Hyperping healthchecks that are down, new or late. For each down one, \
give the last ping and the last 3 runs from get_healthcheck. Ten lines at most." \
  --mcp-config "$MCP_CONFIG" \
  --strict-mcp-config \
  --allowedTools "mcp__hyperping__list_healthchecks" "mcp__hyperping__get_healthcheck"

On the morning the digest failed, list_healthchecks returns something like this (names and IDs shortened):

{
  "counts": { "total": 4, "down": 1, "up": 3, "new": 0 },
  "returned": 4,
  "truncated": false,
  "healthchecks": [
    {
      "uuid": "hc_...",
      "name": "Nightly support digest",
      "state": "down",
      "last_ping_at": "2026-10-07T04:03:12Z",
      "down_since": "2026-10-08T04:25:00Z",
      "schedule": "cron '0 6 * * 1-5' (Europe/Paris)",
      "grace": "25 minutes",
      "ongoing_incident": "hco_..."
    },
    {
      "uuid": "hc_...",
      "name": "Eval suite",
      "state": "up",
      "last_ping_at": "2026-10-07T03:48:40Z",
      "next_ping_due_at": "2026-10-08T03:00:00Z",
      "down_if_no_ping_by": "2026-10-08T06:00:00Z",
      "late": true,
      "schedule": "cron '0 5 * * *' (Europe/Paris)",
      "grace": "3 hours"
    }
  ]
}

The brief can then say: the support digest has been down since 6:25, its last good run was yesterday at 6:03, and the eval suite is 2 hours 15 minutes late but has until 8:00 before it alerts. Send the text wherever your team reads in the morning, and give the brief its own healthcheck: it is a scheduled agent too.

The --allowedTools list lets the two read tools run without a prompt, and --strict-mcp-config keeps other configured servers out of this run. With a read-only key, write tools are refused by the server anyway.

An agent that checks its upstream job first

The same tools stop an agent from working on stale input. The digest reads an export that another job writes at 5:30. Add one line at the top of its prompt:

Before anything else, call get_healthcheck for "Support export". If its state is not up, or its last_ping_at is earlier than today 03:30 UTC, stop and say the export did not run.

The digest then fails its own output check, its healthcheck goes down, and the alert names the real cause instead of an empty summary. The export has its own healthcheck, so both alerts arrive, in order.

The MCP server allows 60 tool calls a minute and 600 an hour per project. A brief that reads a few dozen healthchecks once a day is far from that.

Where to start

Pick the scheduled agent whose silence would cost you the most, often the eval run before a release or the job that feeds a customer-facing report. Give it a healthcheck in cron mode with its own schedule and timezone, wrap the run in /start and a checked success ping, and let a week of durations tell you the right grace period. Then add the morning brief.

For the plain cron jobs around your agents, the same approach is in the guides for Node.js cron jobs, GitHub Actions scheduled workflows and Vercel Cron, and Hyperping for AI companies covers the rest of the stack: inference endpoints, model providers and status pages.