Most infrastructure fails visibly to an outside monitor: the request times out, the port refuses, the certificate is wrong. A VPN is the exception by design. WireGuard listens on UDP and answers nothing that isn't a valid handshake. Tailscale is peer-to-peer with no server to probe. The tunnel exists precisely so that nothing on the public internet can tell whether it's there. Which means when it dies, nothing tells you. You find out when you try to reach the home server from a hotel, or when the office subnet is unreachable on Monday morning and the router rebooted on Saturday without the route.

This guide monitors a tunnel from the only place that can see it: inside. A short script on one peer pings across the tunnel, reads WireGuard's handshake ages or Tailscale's online state and key expiry, probes the internal services the tunnel exists to reach, and pings a CronAlert heartbeat only when everything passes. It also covers the two external checks that still work, and the placement question that decides whether the heartbeat can report the failure it detects.

How VPNs die quietly

  • Tailscale node key expired. Node keys expire after 180 days by default. The node drops off the tailnet; the daemon, the machine, and every service on it are fine. Nothing alerts. This is the single most common "the VPN broke" for servers nobody logs into.
  • Peer stopped handshaking. The other side changed IP without a DNS update, a firewall rule dropped UDP after a router firmware update, or the peer's config was regenerated with a new key. wg show reports a handshake from hours ago and no traffic.
  • Subnet router or exit node came back without its role. The machine rebooted, tailscaled started, but the route advertisement wasn't persisted or IP forwarding wasn't re-enabled. The node is online; the LAN behind it is unreachable.
  • The control server fell over. If you run Headscale, its being down means no new connections, no key renewals, and no DNS for the tailnet. Existing tunnels limp on from cached state until they don't.
  • The VPN host is simply off. Power, a kernel panic, a provider outage. The only failure on this list an external check can see.

What can still be checked from outside

Two things, and both are worth having. If the WireGuard server or a Tailscale subnet router has a public address, a TCP port check on its SSH port tells you the host is up and reachable, which separates "the box is dead" from "the tunnel is broken" in every alert that follows. Don't try to check the WireGuard UDP port itself; a TCP probe to UDP 51820 times out whether the tunnel is healthy or not, and there's no ICMP option either.

If you self-host the control plane with Headscale, monitor it like any web service: an HTTPS check on its public hostname's /health endpoint, expecting a 200, from the free plan. Tailscale's hosted control plane needs no monitor from you (its status page is status.tailscale.com), and an outage there leaves established tunnels working from cached keys; only new connections and rotations fail.

The heartbeat script inside the tunnel

A heartbeat monitor inverts the check: CronAlert gives you a URL, a cron job on a peer runs every minute, and it requests that URL only if every check passed. Silence, from any cause, becomes the alert. Create the heartbeat monitor with a 1-minute expected interval (it alerts after two minutes of silence with the default grace), then install this on a peer:

#!/usr/bin/env bash
# /usr/local/bin/vpn-health.sh — run from cron every minute on a tunnel peer
set -uo pipefail
HEARTBEAT="https://cronalert.com/api/heartbeat/<token>"
PEER_TUNNEL_IP="10.0.0.1"          # the other side's tunnel address (WireGuard) or Tailscale IP
INTERNAL_URLS="http://10.0.0.5:8123/ http://10.0.0.6:9000/"   # services the tunnel exists to reach
fail=0

# 1. Can we cross the tunnel at all?
ping -c1 -W3 "$PEER_TUNNEL_IP" >/dev/null 2>&1 || { echo "no reply from $PEER_TUNNEL_IP through the tunnel"; fail=1; }

# 2a. WireGuard: newest handshake must be recent (needs PersistentKeepalive on the peer)
if command -v wg >/dev/null && wg show wg0 >/dev/null 2>&1; then
  latest=$(wg show wg0 latest-handshakes | awk '{print $2}' | sort -n | tail -1)
  age=$(( $(date +%s) - ${latest:-0} ))
  [ "$age" -gt 180 ] && { echo "wg0 last handshake ${age}s ago"; fail=1; }
fi

# 2b. Tailscale: this node is online and its key isn't about to expire
if command -v tailscale >/dev/null; then
  status=$(tailscale status --json 2>/dev/null)
  echo "$status" | jq -e '.Self.Online' >/dev/null || { echo "tailscale reports this node offline"; fail=1; }
  expiry=$(echo "$status" | jq -r '.Self.KeyExpiry // empty')
  if [ -n "$expiry" ]; then
    left=$(( $(date -d "$expiry" +%s) - $(date +%s) ))
    [ "$left" -lt 604800 ] && { echo "tailscale key expires in $((left/86400)) days"; fail=1; }
  fi
fi

# 3. The services behind the tunnel answer
for url in $INTERNAL_URLS; do
  curl -fsS -m 5 -o /dev/null "$url" || { echo "$url unreachable"; fail=1; }
done

# Ping only when everything passed; silence is the alert
[ "$fail" -eq 0 ] && curl -fsS -m 10 -X POST "$HEARTBEAT" >/dev/null
exit 0
# /etc/cron.d/vpn-health
* * * * * root /usr/local/bin/vpn-health.sh 2>&1 | logger -t vpn-health

Notes on each check. The cross-tunnel ping is the one that matters most; everything else explains why it failed. For Tailscale, tailscale ping is an alternative that works even where ICMP is filtered. The WireGuard handshake check only makes sense with PersistentKeepalive = 25 set on at least one side; WireGuard is silent when idle, so without keepalives an old handshake is normal, not a fault. The Tailscale key check gives you a week's warning on the failure that otherwise gives none; if these are servers, consider disabling key expiry for them in the admin console and keep the check as a backstop. And the internal-service probes are what turn "the VPN is up" into "the thing I use the VPN for is up," which is the question you actually have at the hotel. The internal tools guide covers what to probe on each.

When the heartbeat goes silent, journalctl -t vpn-health -n 20 on the peer names the check that failed. Heartbeat monitors are on the Pro plan and above.

Where to run it, and why two is better

The script has to reach CronAlert when the tunnel is down, so run it on a peer whose path to the internet doesn't go through the tunnel: the VPN server itself, a small VPS, a Raspberry Pi at home on the LAN side. A laptop that routes everything through an exit node is the wrong place; when the tunnel fails, so does its heartbeat, which still produces an alert but not a diagnostic one.

Better still, run it on both ends with two heartbeat monitors. Then the alerts read like a table:

Server host TCP 22Server-side heartbeatClient-side heartbeatMost likely
DownSilentSilentThe VPN host is off or unreachable. Start with power and the provider.
UpSilentSilentThe tunnel itself: handshakes failing both ways. Check firewall rules for UDP, keys, endpoint DNS.
UpFineSilentThe client side lost its route or its internet; the server sees no problem.
UpFineFine, but with internal probe failures in the logThe tunnel is healthy; a service behind it is down. Not a VPN problem.
UpSilent every ~6 monthsFineKey expiry. The script warned a week early; renew or disable expiry for that node.

Subnet routers and exit nodes

A subnet router extends the tailnet to a whole LAN; when it reboots and comes back without the advertised route (or with IP forwarding off), the node is online and every internal check fails. The script above catches that through the internal-service probes, since they target LAN addresses through the router. For an exit node, add one line that confirms traffic actually egresses through it: from a client using the exit node, curl -s https://ifconfig.me should return the exit node's public IP; compare it to the expected value and fail if it doesn't match. That catches the surprisingly common state where the exit node is online but the client fell back to its own connection.

Frequently asked questions

Can CronAlert check the WireGuard port?

No; it's UDP and answers nothing but valid handshakes. Check the host's SSH port from outside and the tunnel from inside.

Why did my Tailscale node vanish?

Usually key expiry, 180 days by default. The script fails a week ahead; or disable expiry for servers.

Where does the script run?

On a peer with internet access outside the tunnel. On both ends if you want to know which side broke.

Which plan?

The Headscale health check is free. Heartbeats and the TCP check on the host are Pro at $5/mo.

What about OpenVPN in TCP mode?

That port can be checked directly with a TCP monitor. The inside-out heartbeat still tells you more.

Watch the tunnel from where it can be seen

You can't monitor a VPN from the internet, and you don't need to. One script on a peer, pinging across the tunnel and reading the daemon's own view of itself, turns every silent failure on the list above into an alert within two minutes, with the reason waiting in the journal. Create an account, upgrade to Pro, and add the heartbeat before the next key expires. Related reading: monitoring a VPS end to end for the host underneath, monitoring a Proxmox host and monitoring a NAS for what usually sits behind the tunnel, monitoring self-hosted apps, and choosing a heartbeat grace period.