Home Assistant has a distinctive way of being broken: everything about it looks fine. The dashboard loads, the history graphs draw, the logbook scrolls, and the lights don't turn on because the Zigbee coordinator fell out of its USB port three hours ago and forty devices have been "unavailable" since. Or the MQTT broker add-on crashed. Or the cloud integration's token expired and the thermostat has been showing yesterday's temperature all day. None of these take Home Assistant down. All of them take your home down.
The other failure mode is the ordinary one: the host lost power, an update hung on a migration, the recorder database filled the disk. That one at least looks broken, but only if you're looking, and nobody opens the dashboard at 3 AM to check. This guide sets up both kinds of alert: an external check on the instance itself, and an automation inside Home Assistant that acts as a heartbeat and goes quiet the moment the things you care about do.
From outside: the API root, with a token
If your instance has a public hostname, either a Nabu Casa remote URL (https://<id>.ui.nabu.casa) or your own domain through a reverse proxy, the best endpoint to monitor is the REST API root, /api/. It requires authentication and answers with a tiny JSON message, which makes it a better liveness signal than the frontend: it proves the core is up and serving API requests, not merely that a cached HTML shell loaded.
- In Home Assistant, open your Profile, then Security, and create a Long-lived access token named "CronAlert". Copy it; it's shown once.
- Create an HTTP monitor on
https://your-hostname/api/with method GET and expected status 200. - Add a custom header:
Authorization: Bearer <your token>. Custom headers are available on every plan. - On Pro, add a keyword check for
API running., the body Home Assistant returns when healthy, so a proxy error page can't pass as up.
If you'd rather not store a token in a monitor, skip the header and set the expected status to 401 instead. Home Assistant answering "unauthorized" is still Home Assistant answering; a timeout or a 502 from the proxy is not. It's a weaker check (it passes even if the core can't reach its database) but it costs nothing and stores nothing.
Two notes for reverse-proxy setups. Home Assistant refuses proxied requests unless the proxy is listed in trusted_proxies under the http: section of configuration.yaml with use_x_forwarded_for: true; if the monitor sees a 400 from a fresh proxy, that's usually why. And because the monitor hits your public hostname, it also tracks the certificate, which catches a stalled Let's Encrypt renewal before the companion app starts refusing to connect. The reverse proxy guide covers the rest of that layer.
If the instance is LAN-only with no public hostname, that's fine; skip this layer. The heartbeat below covers it without exposing anything.
From inside: an automation as a heartbeat
A heartbeat monitor gives you a URL and expects a request on a schedule; silence becomes the alert. Home Assistant can make that request itself, from an automation, and the trick that makes this valuable is the condition: the automation only pings when your critical entities are alive. Create a heartbeat monitor in CronAlert with a 5-minute expected interval (it alerts after 10 minutes of silence with the default grace), then add to configuration.yaml:
rest_command:
cronalert_heartbeat:
url: "https://cronalert.com/api/heartbeat/<token>"
method: POST
timeout: 10 Restart (or reload the REST commands), then create the automation. In YAML, using the current trigger/condition/action syntax:
alias: CronAlert heartbeat
description: Ping CronAlert every 5 minutes while the critical devices are alive
triggers:
- trigger: time_pattern
minutes: "/5"
conditions:
# Zigbee coordinator reachable (Zigbee2MQTT bridge; ZHA exposes similar sensors)
- condition: state
entity_id: binary_sensor.zigbee2mqtt_bridge_connection_state
state: "on"
# A sensor that should always be reporting is not unavailable
- condition: not
conditions:
- condition: state
entity_id: sensor.living_room_temperature
state: unavailable
# ...and that it reported recently (catches a sensor stuck on a stale value)
- condition: template
value_template: "{{ (now() - states.sensor.living_room_temperature.last_updated).total_seconds() < 1800 }}"
actions:
- action: rest_command.cronalert_heartbeat
mode: single Older installations use trigger:/platform:, condition:, and action:/service: keys; the automation editor accepts either. Choose conditions for the three or four things whose failure you'd actually want to be woken for: the Zigbee or Z-Wave coordinator's connection sensor, the MQTT broker (the mqtt integration exposes a connection sensor, or condition on any MQTT-sourced entity not being unavailable), a presence or security sensor, the integration behind your heating. Avoid conditions on things that are legitimately offline sometimes, such as a phone's battery sensor; every condition is a promise that its failure deserves an alert.
The result covers both failure modes with one monitor. If the host dies, the automation never runs and the pings stop. If Home Assistant is running but a coordinator dropped, the condition fails and the pings stop. Either way, CronAlert alerts you within about ten minutes, and the automation's trace in Settings → Automations shows which condition blocked it.
What to alert on, and what the alert should say
Give the heartbeat monitor a name that tells you what silence means ("Home Assistant: core + Zigbee + MQTT"), because that's what the alert will say. If you want separate alerts for separate subsystems, make two automations with two heartbeat monitors: one unconditional (the instance is alive) and one conditional (the devices are alive). Then "both silent" is the host and "only the second silent" is a device problem, and you know which before you open anything.
Send the alert somewhere that doesn't depend on the same house. If your Home Assistant notifies you through the companion app and the instance is what's down, the notification never leaves. CronAlert's channels are independent of it: push to your phone with a Do Not Disturb exception, Discord, email, or Telegram. Set a maintenance window over the update slot if you update on a schedule, so a planned restart doesn't page you.
Failures this catches
- Host down or stuck on an update. Both monitors go silent. HA OS updates that hang on a database migration are the common one.
- Recorder filled the disk. Home Assistant often stays up but stops recording and eventually fails to write anything; the external check may still pass, the heartbeat usually stops once the rest_command can't log.
- Zigbee or Z-Wave coordinator disconnected. USB power management, a loose hub, a crashed Zigbee2MQTT add-on. The condition blocks the ping.
- MQTT broker add-on crashed. Every MQTT entity goes unavailable at once; the condition blocks the ping.
- A cloud integration silently stopped updating. The
last_updatedtemplate condition catches an entity frozen on a stale value, which "unavailable" alone misses. - Reverse proxy or certificate broke. The external check fails while the heartbeat keeps going: the instance is fine, the path to it isn't, and you know to look at the proxy.
Frequently asked questions
Which URL do I monitor?
/api/ on your public hostname, with a long-lived token in an Authorization: Bearer header, expecting 200. Or expect 401 without a token as a weaker, credential-free check.
How do I get alerted when devices go unavailable?
A time-pattern automation whose conditions require the critical entities to be available, calling a rest_command that pings a heartbeat. Silence is the alert.
My instance is LAN-only. Does this work?
The heartbeat does, with nothing exposed. For an external check too, Nabu Casa gives you a public hostname.
Which plan?
The token-header HTTP check is free. The keyword assertion and the heartbeat are Pro ($5/mo).
Does this replace Home Assistant's own notifications?
No, it covers what they can't: the case where the instance that would send them is the thing that's down.
Watch the hub, and watch what it's connected to
A smart home fails at the hub rarely and at the edges constantly, and only an alert that fires on both is worth having. One HTTP check tells you the instance is reachable; one automation with the right conditions tells you the house still works. Create an account, add the API check, and upgrade to Pro for the heartbeat before the next time the stick falls out. Related reading: monitoring Docker and self-hosted apps, monitoring IoT devices, monitoring Pi-hole and AdGuard Home for the DNS your smart home depends on, monitoring a NAS, and choosing between HTTP, TCP, and heartbeat monitors.