A heartbeat monitor is the simplest thing in uptime monitoring: your job pings a URL when it finishes, and if the pings stop, you get an alert. The subtlety is entirely in the word "stop." A backup that finished twenty minutes late hasn't stopped. A backup that hasn't run since Tuesday has. Somewhere between those two is a deadline, and the grace period is how you tell CronAlert where to draw it.

This post explains the two numbers that define a heartbeat's deadline, why the default grace is what it is, how to pick a grace for jobs that run every minute, every hour, every night, and every week, where the clock starts, and the handful of mistakes that produce heartbeat alerts for jobs that were fine.

Two numbers define the deadline

Every heartbeat monitor has an expected interval, which is how often your job is supposed to ping (every minute, every hour, every 24 hours), and a grace period, which is how late a ping may be before the monitor is marked down. The deadline is simply:

deadline = time of last ping + expected interval + grace period

CronAlert checks every heartbeat monitor once a minute against that deadline. When the deadline passes with no ping, an incident opens and your alert channels fire. When the next ping arrives, the incident resolves and a recovery notification goes out. Each ping also moves the deadline forward, so an early ping is never a problem; only a late one is.

The expected interval can be anything from your plan's minimum (one minute on Pro and above) up to seven days. The grace can be anything from zero to seven days. With the default grace, the deadlines look like this:

Expected intervalDefault graceAlert fires when silence reaches
1 minute1 minute2 minutes
5 minutes5 minutes10 minutes
15 minutes15 minutes30 minutes
1 hour1 hour2 hours
6 hours1 hour7 hours
24 hours1 hour25 hours
7 days1 hour7 days, 1 hour

Why the default is "one interval, capped at an hour"

For frequent pings, one interval is the natural unit of doubt. A job that runs every minute and misses one ping might just have hit a slow network; missing two in a row means something is wrong. So a 1-minute heartbeat alerts after two minutes, a 5-minute one after ten, and you never get paged for a single late packet.

Scaling that rule up breaks down. If the grace for a daily job were also one interval, a backup that died on Monday night wouldn't alert until Wednesday night, 48 hours after the last success and a full day after you could have fixed it. Capping the default grace at an hour means a nightly job alerts around 25 hours after its last ping: soon enough to matter, late enough to absorb a slow run. The cap is a default, not a limit; the whole point of the setting is that your job's real behavior should decide it.

Choosing a grace from your job's real behavior

The key observation: because the ping fires at the end of the job, the gap between two consecutive pings isn't the schedule period. It's the schedule period plus however much longer this run took than the last one. A nightly job that took 30 minutes on Monday and 90 minutes on Tuesday pinged 25 hours apart, even though it started exactly on schedule both nights. The grace period exists to absorb that difference plus a few minutes of cron start jitter. So the rule is:

grace ≥ (worst run time − typical run time) + start jitter

and no larger than that, or the alert arrives too late to be useful. Applied to the common schedules:

  • Every minute to every 15 minutes (queue workers, sync loops, the host-health scripts in our VPS and Proxmox guides): keep the default of one interval. These jobs are short and regular, and one missed tick is the signal you want.
  • Hourly (report generation, cache warming, log shipping): the default 1-hour grace is right unless the job's run time varies by more than a few minutes; if it does, that variance is itself worth investigating.
  • Nightly (database backups, ETL, invoice runs): look at the longest run in your logs. A 2 AM dump that normally takes 40 minutes but has taken two hours on a busy month wants a grace of two to three hours. You're paged by 5 or 6 AM instead of at 2 AM the next night, and a normal slow run doesn't wake you.
  • Weekly or monthly (restore tests, certificate renewals, archive rotations): a grace of 12 to 24 hours. At these frequencies the risk isn't a slow run, it's the schedule silently drifting or being disabled, and a day is a fine resolution for noticing that.

If a job's run time varies so much that the grace would have to be longer than the alert is useful, the heartbeat isn't the right shape for it. Split it into a start ping and a finish ping on two monitors so you can alert on "didn't start" and "started but never finished" separately; the long-running batch job guide covers that pattern.

Where the clock starts

Three details about the clock that people ask about:

  • A new monitor counts from its creation. You have a full interval plus grace to paste the ping URL into your job and let it run once. A daily heartbeat created at 10 AM won't alert until 11 AM the next day if no ping ever arrives, which is exactly when it should.
  • CronAlert doesn't know your schedule, only your spacing. It has no idea the job runs at 2 AM; it knows pings arrive about 24 hours apart. That's a feature: a manual run at 3 PM simply moves the deadline to 3 PM tomorrow plus grace, the scheduled 2 AM run pings well before that, and the deadline moves again. Early pings never cause alerts.
  • Recovery is automatic. When a ping arrives during an incident, the incident resolves and a recovery alert goes out with the outage duration. You don't acknowledge heartbeats; you fix the job.

For planned silence, such as pausing backups during a migration, set a maintenance window on the heartbeat monitor rather than widening its grace. The grace should describe the job's normal behavior; the window describes the exception.

Five mistakes that cause false heartbeat alerts

  1. Setting the interval to the job's duration instead of its schedule. A job that runs nightly and takes 40 minutes has an expected interval of 24 hours, not 40 minutes. The interval is the spacing between pings.
  2. A grace of zero on a cron job. Cron starts jobs within the scheduled minute, not on the second, and the ping itself takes a moment to arrive. Zero grace only makes sense for pings driven by a tight loop where any gap is a real fault.
  3. Pinging at the start of the job. Then a job that starts and crashes looks healthy. Ping on success, at the end, and only after verifying the output: backup.sh && curl -fsS "$HEARTBEAT". The heartbeat guide covers the exit-code traps.
  4. A grace longer than the interval on a frequent job. A 5-minute heartbeat with a 1-hour grace hides an eleven-minute outage entirely and a two-hour outage for most of its length. For jobs under an hour, one interval is almost always the right grace.
  5. Forgetting the grace when you change the schedule. Moving a job from nightly to weekly without changing the monitor's interval produces an alert every day at the old deadline. Update the interval and grace together, in the same change as the crontab.

Setting the interval and grace

In the dashboard, create a monitor and choose the Heartbeat tab. Pick an Expected Interval from one minute to seven days and a Grace Period, or leave the grace on Auto to use the default; the Auto option shows what it resolves to for the interval you picked. Both can be changed later from the monitor's edit page, and the monitor's detail page shows the effective grace.

Through the REST API, pass the expected interval and grace in seconds when creating a heartbeat monitor:

curl -X POST https://cronalert.com/api/v1/monitors \
  -H "Authorization: Bearer $CRONALERT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"type":"heartbeat","name":"Nightly DB backup","checkInterval":86400,"gracePeriod":7200}'

The response includes the monitor's ping URL. To change the grace later, PUT the monitor with a new gracePeriod; sending null returns it to the default. For HTTP and TCP monitors the interval is set by your plan and checkInterval is ignored.

From an AI assistant connected to the CronAlert MCP server (version 1.5.0 or later), plain language is enough: "create a heartbeat called Nightly DB backup that expects a ping every 24 hours with a 2-hour grace, and give me the URL." The assistant fills in the same two fields and hands back the ping URL.

Heartbeat monitors are on the Pro plan and above.

Frequently asked questions

What is the grace period?

Extra time a ping may be late before the monitor goes down. Deadline = last ping + expected interval + grace.

What's the default?

One interval, capped at one hour. A 1-minute heartbeat alerts after 2 minutes; a daily one after 25 hours.

How do I pick a grace for a nightly job?

Worst run time minus typical run time, plus a few minutes. Usually two to three hours for a backup.

Will a new monitor alert before I've added the ping?

No. It counts from creation, so you get interval plus grace to wire it up.

Can I set the grace to zero?

Yes, but only for pings from a tight loop. Cron jobs need at least a minute or two for start jitter.

Does CronAlert know my job runs at 2 AM?

No, only that pings are about 24 hours apart. Manual or early runs never cause alerts; they just move the deadline.

Set it once, from the logs

The right grace period is in your job's history: the longest it has ever legitimately taken, minus the usual, plus a little slack. Set the expected interval to the schedule, the grace to that number, and the heartbeat will stay quiet through every slow night and speak up the first morning the job didn't run. Create an account, upgrade to Pro, and give your nightly backup a deadline. Related reading: cron job heartbeat monitoring, monitoring scheduled database backups, monitoring long-running batch jobs, monitoring systemd timers, monitoring Windows Task Scheduler tasks, and testing your alerts before a real miss does.