A self-hosted forge is a small platform pretending to be one app. Gitea or Forgejo serves a web UI and an API, speaks Git over HTTPS and over SSH on a separate port, hands jobs to Actions runners that are separate processes on possibly separate machines, pulls mirrors from upstreams on a schedule, hosts a container registry, and depends on a nightly gitea dump that nobody checks. Each of those fails on its own, and most of them fail quietly: the web UI loads fine while the runner has been offline since Tuesday and every push has been sitting at "waiting" in a yellow circle.

This guide monitors the layers separately, so an alert tells you which one broke: the forge's own health endpoint, a check that proves Git is actually serving, the SSH port, a heartbeat that the runner itself produces, mirror freshness, the registry, and the backup.

The web layer: /api/healthz

Both Gitea and Forgejo expose /api/healthz without authentication. It isn't a static 200: the forge pings its database and cache and returns 200 with "status":"pass" when they answer, or 503 with "status":"fail" and the failing check named when they don't. That makes it the right first monitor: an HTTP check on https://git.example.com/api/healthz expecting 200 already separates "nginx is up" from "the forge can reach Postgres." Create it on the free plan; it also tracks the certificate on your hostname. On Pro, add a keyword check for pass so a reverse proxy's own 200 error page can't impersonate a healthy forge.

If the forge requires sign-in to view anything (REQUIRE_SIGNIN_VIEW), /api/healthz still answers; /api/v1/version, another unauthenticated endpoint that returns the running version, may not. Prefer healthz.

Prove Git is serving, not just the UI

The health endpoint can pass while Git itself is broken: a corrupt repository root permission, a hook failing, an LFS store unmounted. The cheapest proof that Git works is the smart-HTTP handshake every clone starts with. For any public repository, GET https://git.example.com/owner/repo.git/info/refs?service=git-upload-pack returns 200 and a list of refs with no authentication; create a tiny public repo named canary for exactly this and monitor that URL expecting 200. If every repository is private, the same URL returns 401, and a monitor expecting 401 still proves the Git HTTP backend answered. On Pro, a keyword check for refs/heads/main (or your default branch) confirms it served real refs.

SSH: a TCP check on the Git port

Most developers push over SSH, and the SSH port is independent of the web layer: Gitea's built-in SSH server or the host's sshd with the forge's shell, on 22 or, very commonly in Docker setups, 2222. A TCP port monitor on git.example.com:2222 (or the bare IP, which TCP checks can reach) alerts when nothing is listening: the container didn't publish the port after a recreate, a firewall rule changed, sshd failed to start. It doesn't authenticate or run Git, so it proves reachability, not that keys work; pair it with the smart-HTTP check above and you have both halves. TCP monitors are on Pro.

Actions runners: make the runner prove it's alive

This is the layer that fails most quietly. A runner (act_runner for Gitea, forgejo-runner for Forgejo) is a separate process registered with the forge; when it dies or loses its registration, the forge shows it as offline in the runners page and every job queues forever with no notification to anyone. Rather than scrape the runners page, make the runner produce a signal. Create a repository, add a workflow on the default branch:

# .gitea/workflows/heartbeat.yml  (Forgejo: .forgejo/workflows/heartbeat.yml)
name: runner heartbeat
on:
  schedule:
    - cron: "*/15 * * * *"
jobs:
  ping:
    runs-on: ubuntu-latest        # the label your production jobs use
    steps:
      - run: curl -fsS -m 10 -X POST "${{ secrets.CRONALERT_HEARTBEAT }}"

Create a heartbeat monitor with a 15-minute expected interval and put its URL in the repository's secrets. Now a ping arrives only if the forge's scheduler fired, a runner with that label picked the job up, the job container started, and it could reach the internet, which is every link in the chain a real pipeline needs. When the runner is offline, stuck on a hung job, or re-registered with different labels, the pings stop and you hear about it within about half an hour. If you run several runners with different labels, one tiny workflow per label tells you which pool is dead. The same pattern monitors GitHub Actions scheduled workflows; the forge version has the extra virtue of testing infrastructure you own.

If your CI is external (Woodpecker or Drone triggered by forge webhooks), add an HTTP monitor on the CI server too: Woodpecker serves /healthz on its web port, and a scheduled pipeline there can carry its own heartbeat the same way.

Mirrors: freshness through the API

Pull mirrors sync on an interval and stop silently when the upstream's credentials expire or the sync job fails. The repository API exposes when a mirror last synced. A small script, run hourly from the host (or anywhere with a read-only API token), checks every mirror and pings a heartbeat only when all are fresh:

#!/usr/bin/env bash
# /usr/local/bin/mirror-health.sh — hourly; pings only if every mirror synced within MAX_AGE
set -uo pipefail
FORGE="https://git.example.com"; TOKEN="<read-only api token>"
HEARTBEAT="https://cronalert.com/api/heartbeat/<token>"
MIRRORS="infra/upstream-tool team/vendor-lib"; MAX_AGE=$((26*3600)); fail=0
for repo in $MIRRORS; do
  updated=$(curl -fsS -H "Authorization: token $TOKEN" "$FORGE/api/v1/repos/$repo" | python3 -c 'import sys,json,datetime; t=json.load(sys.stdin)["mirror_updated"]; print(int(datetime.datetime.now(datetime.timezone.utc).timestamp()-datetime.datetime.fromisoformat(t.replace("Z","+00:00")).timestamp()))') \
    || { echo "$repo: api error"; fail=1; continue; }
  [ "$updated" -lt "$MAX_AGE" ] || { echo "$repo: last sync ${updated}s ago"; fail=1; }
done
[ "$fail" -eq 0 ] && curl -fsS -m 10 -X POST "$HEARTBEAT" >/dev/null
exit 0

Set the heartbeat's expected interval to one hour and MAX_AGE to the mirror interval plus slack. A stale mirror then surfaces as a heartbeat that went quiet, with the repository named in the journal.

Registry and backups

If you use the forge's package registry for containers, the OCI endpoint https://git.example.com/v2/ answers an unauthenticated request with 401 and a WWW-Authenticate challenge, which is the registry saying it's alive; a monitor expecting 401 catches it being down without needing credentials. And the backup: gitea dump (or forgejo dump) from a nightly cron produces a zip of the database, repositories, and config, and it's only a backup if it ran and completed. End the cron line with a heartbeat ping on success (gitea dump -f /backups/gitea-$(date +%F).zip && curl -fsS …), set the expected interval to 24 hours, and give the grace an hour or two for a growing repository store. The backup monitoring guide covers verifying the dump before pinging.

Reading the alerts

healthzSmart-HTTPSSH TCPRunner heartbeatMost likely
DownDownDownSilentHost or network. Everything is unreachable.
503DownUpSilentDatabase or cache; the forge is up but can't serve. healthz names the check.
UpDownUpFineGit HTTP backend, LFS, or repository storage.
UpUpDownFineSSH port only: container port mapping, sshd, firewall.
UpUpUpSilentThe runner. Jobs are queuing.

Frequently asked questions

Which URL do I monitor?

/api/healthz, expecting 200; it checks the database and cache. Add a smart-HTTP check on a public canary repo to prove Git serves.

How do I know a runner is offline?

A 15-minute scheduled workflow that pings a heartbeat. No ping, no runner.

How do I catch a stale mirror?

Compare mirror_updated from the repo API to now and ping a heartbeat only when all mirrors are fresh.

Which plan?

HTTP checks are free. TCP on the SSH port, keywords, and heartbeats are Pro ($5/mo).

Does this apply to Gogs?

Partly: the SSH and backup patterns do, but Gogs lacks /api/healthz and Actions; use the smart-HTTP check as its primary monitor.

A forge is a platform; monitor it like one

Five small monitors cover the five ways a forge fails: the health endpoint, a clone handshake, the SSH port, a workflow that pings only when a runner ran it, and a heartbeat on the nightly dump. The next time a push sits at "waiting," you'll know before the developer does. Create a free account, add the health and clone checks, and upgrade to Pro for the runner heartbeat. Related reading: monitoring Docker and self-hosted apps, uptime monitoring in CI/CD pipelines, monitoring GitHub Actions scheduled workflows, monitoring Nextcloud and Immich, and choosing between HTTP, TCP, and heartbeat monitors.