A NAS is designed to survive failure quietly. A disk dies and the array keeps serving. A backup job fails and the files from last week are still there. The volume creeps toward full and nothing breaks until the day it does. That resilience is the point of owning one, and it's also why NAS failures are discovered late: the box is still answering, so nobody looks at it. Synology, TrueNAS, and Unraid all have their own notification systems, and all of them share one weakness: they run on the device that might be the thing that's broken, and they send nothing at all when the device is off.
This guide sets up two layers for a NAS. External checks on the services you deliberately publish, which catch the NAS being unreachable. And a short scheduled script on the NAS itself that checks array health, capacity, and the services that matter, then pings a heartbeat only when everything passes, which catches every degraded-but-serving state without exposing a single port. It covers Synology DSM, TrueNAS, and Unraid specifically.
Layer 1: from outside, monitor what you publish, not the admin UI
The tempting first monitor is an HTTPS check on the admin interface: DSM on port 5001, the TrueNAS or Unraid UI on 443. Two problems. All three ship with a self-signed certificate, so an external HTTPS check fails the TLS handshake and reports a certificate error rather than a status. You can fix that (DSM issues Let's Encrypt certificates from Control Panel → Security → Certificate, TrueNAS has ACME support, Unraid offers a certificate for its myunraid.net hostname), but the bigger problem is that the admin interface probably shouldn't be reachable from the internet in the first place. Don't open it to monitor it.
Instead, monitor the things you already publish. Most NAS owners run a handful of services through a reverse proxy with a real hostname: Synology Photos or Drive, Nextcloud, a Plex or Jellyfin server reading from the array, a Vaultwarden container on Unraid. One HTTP monitor per public hostname, expecting a 200, from CronAlert's free plan every three minutes. If the NAS loses power, its network, or its reverse proxy, every one of those goes red together, which tells you it's the box. If one goes red alone, it's that service. The reverse proxy guide covers reading the difference between a 502 and a timeout.
If the NAS has a public IP or a DDNS hostname and you've deliberately exposed a non-HTTP port (SFTP for a partner drop, say), a TCP port check on that port confirms reachability. Don't do this for SMB on 445; SMB should never face the internet, monitored or not. And if the NAS lives entirely behind NAT with nothing published, skip this layer; the next one covers it.
Layer 2: a heartbeat from inside the NAS
Every failure that makes a NAS owner's week bad happens while the NAS is up. A heartbeat monitor catches these: CronAlert gives you a URL, a scheduled script on the NAS checks what matters and requests that URL only if every check passes, and if the requests stop, an incident opens. The NAS being offline and the NAS being degraded produce the same silence, which is what you want. The expected interval and grace period set how quickly silence becomes an alert: a 5-minute schedule with the default grace alerts after 10 minutes.
What the script should check is the same on every platform: is the array healthy, is the main volume below a capacity threshold, are the services people depend on running, and did the last backup succeed recently. How you ask each question differs per platform.
Synology DSM
DSM uses Linux md for SHR and RAID volumes, so array health is in /proc/mdstat, where a healthy array shows every member as U and a degraded one shows an underscore. Create the script as a scheduled task: Control Panel → Task Scheduler → Create → Scheduled Task → User-defined script, run as root, schedule every 5 minutes, and paste:
# Synology DSM — Task Scheduler user-defined script, every 5 minutes as root
HEARTBEAT="https://cronalert.com/api/heartbeat/<token>"
fail=0
# Array: any "_" in the member map means a degraded md array
grep -E '^\s+[0-9]+ blocks.*\[[U_]+\]' /proc/mdstat | grep -q '_' && { echo "array degraded"; fail=1; }
# Capacity on the main volume
use=$(df --output=pcent /volume1 | tail -1 | tr -dc '0-9')
[ "$use" -ge 90 ] && { echo "volume1 at ${use}%"; fail=1; }
# Services that must be running (SMB, Docker/Container Manager if you use it)
for svc in pkgctl-SMBService pkgctl-ContainerManager; do
systemctl is-active --quiet "$svc" || { echo "$svc not active"; fail=1; }
done
# Ping only when everything passed
[ "$fail" -eq 0 ] && curl -fsS -m 10 -X POST "$HEARTBEAT" >/dev/null
exit 0 Service unit names vary by DSM version and installed packages; run systemctl list-units 'pkgctl-*' over SSH once to get the names for your box, and drop the line if you'd rather not depend on them. Task Scheduler can also email you the script's output when it fails, which is a convenient way to see which line tripped.
TrueNAS (SCALE and CORE)
TrueNAS is ZFS, so zpool status -x answers the array question in one line: all pools are healthy, or the name of the pool that isn't. Add the script under System Settings → Advanced → Cron Jobs (SCALE) or Tasks → Cron Jobs (CORE), run as root every 5 minutes:
# TrueNAS — cron job, every 5 minutes as root
HEARTBEAT="https://cronalert.com/api/heartbeat/<token>"
POOL=tank
fail=0
zpool status -x | grep -q 'all pools are healthy' || { echo "zfs: $(zpool status -x | head -1)"; fail=1; }
# Capacity: ZFS pools degrade in performance past ~80%
cap=$(zpool list -H -o capacity "$POOL" | tr -dc '0-9')
[ "$cap" -ge 85 ] && { echo "$POOL at ${cap}%"; fail=1; }
# SMB is serving
systemctl is-active --quiet smbd 2>/dev/null || service samba_server status >/dev/null 2>&1 || { echo "smb not running"; fail=1; }
[ "$fail" -eq 0 ] && curl -fsS -m 10 -X POST "$HEARTBEAT" >/dev/null
exit 0 The SMB line tries the SCALE (Linux) form first and the CORE (FreeBSD) form second, so the same script works on either. TrueNAS's own alert system is good; the heartbeat doesn't replace it, it covers the case where the alert system's host is the thing that's gone.
Unraid
Unraid's array state comes from mdcmd status, which prints mdState=STARTED when the array is up and one rdevStatus.N= line per slot, where DISK_OK is healthy and DISK_DSBL is a disabled (failed or red-balled) disk. Install the User Scripts plugin from Community Applications, add a script, and set its schedule to a custom cron of */5 * * * *:
# Unraid — User Scripts plugin, custom schedule */5 * * * *
HEARTBEAT="https://cronalert.com/api/heartbeat/<token>"
fail=0
status=$(mdcmd status)
echo "$status" | grep -q 'mdState=STARTED' || { echo "array not started"; fail=1; }
echo "$status" | grep -q 'rdevStatus.*=DISK_DSBL' && { echo "a disk is disabled"; fail=1; }
# Capacity on the user share pool
use=$(df --output=pcent /mnt/user | tail -1 | tr -dc '0-9')
[ "$use" -ge 90 ] && { echo "/mnt/user at ${use}%"; fail=1; }
# Docker containers that must be running
for c in plex vaultwarden; do
[ "$(docker inspect -f '{{.State.Running}}' "$c" 2>/dev/null)" = true ] || { echo "container $c not running"; fail=1; }
done
[ "$fail" -eq 0 ] && curl -fsS -m 10 -X POST "$HEARTBEAT" >/dev/null
exit 0 Unraid runs a parity check periodically and the array stays fully available during it, so the script doesn't treat a running check as a failure. If you want to know that the scheduled parity check actually ran, the same freshness trick from the Proxmox guide applies: have the check's completion touch a marker file and fail the script when the marker is older than your schedule.
Backups that quietly stopped
A NAS is usually the destination of backups, the source of them, or both, and a backup job that silently stops is the classic NAS failure. If your backups are scripts you control (rsync to a second NAS, restic to cloud storage, a Time Machine target you rotate), end them with a heartbeat ping on success, exactly as in the backup monitoring guide, with the expected interval set to the schedule. If they're driven by a built-in app like Hyper Backup or Cloud Sync that gives you no hook, check the result instead: a line in the health script that fails when the newest file in the destination is older than your schedule allows. Either way, the alert is the same shape as everything else on this page: a heartbeat that stopped.
Disk hibernation, scrubs, and other planned quiet
Two things to tune. If you rely on disk hibernation to save power, know that a per-minute script won't itself spin up data disks (reading /proc/mdstat or zpool status doesn't touch them), but writing its output to a file on a data volume will, and on Synology the system partition lives on the disks regardless. Run the script every 5 to 15 minutes rather than every minute, log to the system journal rather than a volume, and accept that detection takes about twice the interval. And put maintenance windows over DSM update reboots, TrueNAS upgrades, and anything else that takes the NAS down on purpose, so a planned reboot doesn't page you.
Reading the alert
| Published services | NAS heartbeat | Most likely |
|---|---|---|
| All down | Silent | The NAS is off or unreachable: power, network, a failed boot after an update. |
| All fine | Silent | Degraded but serving. The script's output (Task Scheduler email, TrueNAS cron log, User Scripts log) names the check: array, capacity, service, or backup freshness. |
| One down | Fine | That service or its container; the NAS is healthy. |
| All down | Fine | The reverse proxy, DDNS, or your internet connection; the NAS itself is fine. |
Frequently asked questions
Can an external check see a degraded array?
No. A degraded array serves files normally. Only a script on the NAS can see it, which is what the heartbeat is for.
Why does HTTPS on port 5001 fail?
Self-signed certificate. And the admin UI shouldn't face the internet anyway; monitor published services instead.
Will the script stop my disks hibernating?
Not by itself, but logging to a data volume will. Run it every 5 to 15 minutes and log to the journal.
Which plan?
HTTP checks on published services are free. The heartbeat and TCP checks are Pro at $5/mo.
Does this replace DSM or TrueNAS notifications?
No. Keep them for detail. The heartbeat covers the case they can't: the NAS itself is gone.
The box that survives failure still needs someone watching
A NAS hides its failures well, which is exactly why it needs an external observer. One HTTP monitor per published service tells you when it's unreachable; one five-line script and a heartbeat tell you when it's quietly degraded. Create an account, add the monitors for what you publish, and upgrade to Pro for the heartbeat before the next disk goes. Related reading: monitoring a Plex or Jellyfin server, monitoring Docker and self-hosted apps, monitoring a Proxmox host, monitoring scheduled backups, and choosing a heartbeat grace period.