Splunk On-Call (the product formerly known as VictorOps) is where uptime alerts go when "someone should look at this" needs to become "this specific person's phone rings until they acknowledge." It adds on-call rotations, escalation, and a timeline that stitches alerts and chat into one record. CronAlert produces the alerts; this guide wires the two together.

The wiring needs one honest caveat up front. CronAlert's webhook channel sends its own JSON document and Splunk On-Call's REST endpoint expects its own, and they don't match, so a Splunk URL pasted straight into a CronAlert webhook channel produces nothing. The fix is a small relay, and it turns out to be an advantage: the relay is also where deduplication, auto-resolve, severity, and per-team routing live, under your control, in about thirty lines.

Step 1: Create the REST endpoint and a routing key

  1. In Splunk On-Call, open Integrations → 3rd Party Integrations, find REST Generic, and enable it (or open it if it already is).
  2. Copy the Service API Endpoint. It looks like https://alert.victorops.com/integrations/generic/20131114/alert/<api-key>/$routing_key. The $routing_key at the end is a literal placeholder you will replace.
  3. Open Settings → Routing Keys and create one per team that should receive CronAlert alerts, named descriptively: cronalert-prod, cronalert-billing. Map each to the escalation policy for that team. An unmapped key lands alerts in the default team, which is rarely who you want paged.

Treat the endpoint URL as a secret; anyone holding it can open incidents in your rotation.

Step 2: Deploy the relay

Create a Cloudflare Worker (free at alert volumes), set three secrets with wrangler secret put: CRONALERT_WEBHOOK_SECRET (a random string you'll also enter in CronAlert), VICTOROPS_API_KEY (from the endpoint URL), and DEFAULT_ROUTING_KEY. Then deploy:

// CronAlert webhook → Splunk On-Call REST relay
export default {
  async fetch(request, env) {
    if (request.method !== "POST") return new Response("ok");
    const raw = await request.text();
    if (!(await verify(raw, request.headers.get("X-CronAlert-Signature"), env.CRONALERT_WEBHOOK_SECRET))) {
      return new Response("bad signature", { status: 401 });
    }
    const a = JSON.parse(raw);
    const down = a.event === "monitor.down";

    // Optional per-team routing from a name prefix like "[billing] Stripe webhooks"
    const tag = a.monitor.name.match(/^\[([a-z0-9-]+)\]/i)?.[1]?.toLowerCase();
    const routingKey = tag ? `cronalert-${tag}` : env.DEFAULT_ROUTING_KEY;

    const body = {
      message_type: down ? "CRITICAL" : "RECOVERY",
      entity_id: `cronalert:${a.monitor.url}`,            // stable per monitor → one incident, auto-resolved
      entity_display_name: `${a.monitor.name} is ${down ? "DOWN" : "back up"}`,
      state_message: down
        ? `${a.monitor.url} — ${a.check.statusCode ?? a.check.errorMessage ?? "no response"}${a.check.region ? ` (${a.check.region})` : ""}`
        : `${a.monitor.url} recovered after ${minutes(a.incident)} min`,
      state_start_time: Math.floor(new Date(a.incident.startedAt).getTime() / 1000),
      monitoring_tool: "CronAlert",
    };

    const res = await fetch(`https://alert.victorops.com/integrations/generic/20131114/alert/${env.VICTOROPS_API_KEY}/${routingKey}`, {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify(body),
    });
    return new Response(null, { status: res.ok ? 204 : 502 });
  },
};

const minutes = (i) => Math.max(1, Math.round((new Date(i.resolvedAt) - new Date(i.startedAt)) / 60000));

async function verify(raw, header, secret) {
  if (!header || !secret) return false;
  const key = await crypto.subtle.importKey("raw", new TextEncoder().encode(secret), { name: "HMAC", hash: "SHA-256" }, false, ["sign"]);
  const sig = await crypto.subtle.sign("HMAC", key, new TextEncoder().encode(raw));
  const hex = [...new Uint8Array(sig)].map((b) => b.toString(16).padStart(2, "0")).join("");
  return hex.length === header.length && crypto.subtle.timingSafeEqual(new TextEncoder().encode(hex), new TextEncoder().encode(header));
}

What the relay receives from CronAlert is documented in webhook alert integrations: event, monitor.name and monitor.url, incident.startedAt and incident.resolvedAt, and check.statusCode, errorMessage, and region. Returning 502 when Splunk rejects the post makes the failure visible in CronAlert's notification log instead of disappearing.

Step 3: Add the webhook channel in CronAlert

  1. Open Alert Channels, add a channel of type Webhook, and name it for the destination: Splunk On-Call (prod).
  2. Paste the Worker's URL, enter the same signing secret you stored in the Worker, and save. The webhook channel has exactly these two settings; CronAlert always POSTs JSON and signs it when a secret is present.
  3. Use Send test alert. An incident should open in Splunk On-Call within seconds and page the rotation on the routing key.

The channel is team-wide: every monitor in this CronAlert team now pages through it, alongside whatever other channels are enabled (Slack for awareness, email for the record). There is nothing to attach per monitor.

Deduplication and recovery

Splunk On-Call merges alerts sharing an entity_id into one incident and closes it when a RECOVERY arrives for the same id. The relay derives the id from the monitor URL, so each monitor maps to exactly one incident at a time; the down event opens it, the recovered event closes it, and the engineer sees it resolve itself on the timeline with the outage duration in the message. CronAlert already emits a single down and a single recovered event per incident (it does not re-send while a monitor stays down), so there is no alert flood to tame; deduplication here is insurance against a monitor that flaps, where it collapses the pairs into one incident's timeline rather than a fresh page each time. For keeping flapping from happening at all, see reducing false positives.

Routing different monitors to different teams

CronAlert alert channels belong to a team, not to individual monitors, so "attach this channel to these monitors" isn't a setting. Two patterns deliver per-team paging anyway:

  • One CronAlert team per on-call team. CronAlert teams are free to create on every plan (inviting members needs the Team plan). Put the billing monitors in a Billing team with a webhook channel pointed at a Worker configured for cronalert-billing, the platform monitors in a Platform team, and so on. Routing is then visible in the team structure.
  • One team, routing in the relay. The Worker above reads a [tag] prefix from the monitor's name and uses it as the routing key suffix, so [billing] Stripe webhooks pages billing and anything unprefixed pages the default key. One channel, one Worker, routing encoded where you can see it: in the monitor's name.

Severity mapping

Splunk On-Call pages on CRITICAL, posts WARNING to the timeline without paging (by default), and treats INFO as a note. The relay sends CRITICAL for every down event, which is right for production. For staging and internal tools that shouldn't wake anyone, either keep them in a CronAlert team whose Worker sends WARNING, or extend the prefix trick: a [warn] tag that flips message_type. Keep the number of severities small; the alert fatigue guide explains why two levels usually beat five.

Testing the whole path

  1. Splunk side first. curl a CRITICAL payload with a test entity_id straight at the REST endpoint, confirm the page, then post a RECOVERY for the same id and confirm it closes. This proves the routing key and escalation policy.
  2. Relay next. CronAlert's Send test alert on the channel; confirm an incident opens.
  3. Real monitor last. Point a spare monitor at a URL that returns 500, wait for the down event, confirm the page, fix the URL, and confirm the recovery closes the incident. The most common failure is "the test works and the recovery doesn't," because recoveries carry a resolvedAt and no status code; the relay above handles both.

Common pitfalls: leaving the literal $routing_key in a URL (alerts vanish to a nonexistent key), a routing key with no escalation policy (alerts land in the default team), and forgetting that the Worker is itself an endpoint; put an HTTP monitor on it, per monitoring webhook receivers.

Frequently asked questions

Is Splunk On-Call the same as VictorOps?

Yes, renamed after Splunk's 2018 acquisition. The URLs still say victorops.

Is there a native Splunk On-Call channel?

No. The payload shapes differ, so a thirty-line relay does the mapping and adds routing and severity control.

How do alerts deduplicate and resolve?

Stable entity_id per monitor; CRITICAL opens, RECOVERY closes.

Can monitors page different teams?

Yes, via one CronAlert team per on-call team or a name-prefix rule in the relay. Channels are team-wide, not per monitor.

Is the webhook channel free?

Yes, on every plan, as is the Worker at this volume.

Get started

Ten minutes: a routing key, a Worker, a webhook channel, and three tests. The next time a monitor fails, the right rotation gets paged with the URL and the error in the message, and the incident closes itself when the monitor recovers. Create a free CronAlert account and start with the test curl. Related reading: PagerDuty alerts (a native channel, no relay), Opsgenie alerts, webhook alert integrations, on-call for small teams, and incident response for small teams.