How it works¶
You run a monitoring server (Uptime Kuma, Prometheus, whatever) that alerts you when things break. But who alerts you when the monitoring server breaks? gha-deadman runs outside your network, on GitHub's runners, and watches a URL that is only healthy while your watcher is healthy.
The pieces¶
.github/workflows/deadman.ymlruns on a*/10cron on GitHub-hosted runners (free and unlimited on public repositories).scripts/check.shprobesTARGET_URL(2 attempts, 25 s apart, so a single network blip doesn't page you) and talks to the Telegram Bot API directly: no third-party actions, no dependencies beyondcurl,jqandgh.- Stateless by design: the workflow's own run history is the state. A
failing probe exits non-zero, so the previous run's conclusion says whether
the target was already down, the streak of consecutive red runs drives the
re-alert cadence, and the oldest red run's timestamp gives the outage
duration. Nothing is stored anywhere, and the run history doubles as an
outage log. (This also sidesteps a real limitation: the workflow
GITHUB_TOKENcannot write repository Actions variables.) - Alert policy: one message on the up-to-down transition, a reminder every hour while down, one message on recovery with the outage duration. Steady state sends nothing.
- A separate weekly
keepalive.ymlre-enables both workflows through the GitHub API so the schedules survive GitHub's 60-day inactivity auto-disable. It is separate on purpose, so its green runs never pollute the probe's history.
How fast it notices¶
Much slower than the cron suggests. GitHub's scheduler is explicitly
best-effort: it delays and drops scheduled runs, and nothing in the workflow can
make it keep a */10 pace. The scheduled runs of this repository:
| Period | Scheduled runs | Median gap | 90% of gaps under | Longest gap |
|---|---|---|---|---|
| 15 to 25 August 2026 | 437 | 30 min | 51 min | 1 h 50 min |
| 29 August to 30 September 2026 | 216 | 3 h 32 min | 5 h 25 min | 8 h 20 min |
So expect to hear about an outage within a few hours, and in the worst gap measured, more than eight. That is fine for "my monitoring host died" and useless for "my API had a 90-second blip": this is a dead-man's switch, not an uptime SLA monitor. Tightening the cron does not help; the throttling is on GitHub's side.
Because the cadence is unreliable, the reminder interval is measured in elapsed
outage time (REALERT_SECONDS), not in number of runs. There is at most one
reminder per REALERT_SECONDS of outage (one per outage hour by default), sent
by the first run in each new hour. Two reminders can still arrive minutes
apart: a run at minute 119 of an outage sends the first-hour reminder, and a
run at minute 121 sends the second-hour one. When runs are more than an hour
apart, as they have been since late August, every run during an outage sends
one.