ChangeTrace
Incidents

Uptime and outages

How ChangeTrace checks your site is up, what each failure mode means, and why some outages get no ranked causes.

Uptime monitoring runs automatically. There is nothing to switch on and nothing to configure.

Every five minutes, from outside your hosting, ChangeTrace requests your site's homepage. This is independent of the plugin — which matters, because when your site is truly down the plugin cannot tell anyone.

When an outage becomes an incident

  • Two consecutive failures — about ten minutes of real downtime — open a critical incident.
  • One success closes it, and records how long it was down.

The two-failure rule is flap suppression: a single timed-out check during a traffic spike is not an outage, and waking you for it would train you to ignore the alerts.

The four health states

Shown on your Overview. Two of these are routinely confused, and the difference is important.

StateWhat it means
HealthyA heartbeat arrived recently, no open incidents, no error burst
DegradedThe site responds, but there is an open incident or a cluster of errors
Heartbeat MissingNo heartbeat in the last 2 hours — and nothing more than that
DownProven: ChangeTrace could not load your site from outside, twice in a row

Heartbeat Missing does not mean your site is down

It means the plugin has not checked in. Your shop may be serving customers perfectly.

The usual cause is WP-Cron not running — see Before you start. Other causes: the plugin was deactivated, the site is in safe mode after a migration, or outbound requests are being blocked.

Down is the only state that means your site is actually unreachable, and ChangeTrace will only say it after proving it from outside.

Failure modes

When a check fails, ChangeTrace works out how, using the status code, markers in the response body, and a second request for a static WordPress file. That last one is the clever part: if a plain .js file still serves while your homepage 500s, the web server is fine and PHP is broken.

Failure modeWhat you are told
WordPress critical errorWordPress ran and crashed, so a recent change on the site is a credible cause.
Stuck in maintenance modeAn update started and never finished, so the update itself is the cause.
Server errorThe web server or PHP process failed. No site change is implicated.
Gateway or upstream errorA proxy or CDN could not reach the origin. No site change is implicated.
Connection refusedThe host refused connections. This is hosting or networking, not a site change.
DNS failureThe domain did not resolve. Check DNS or domain renewal — not a site change.
TLS/certificate failureThe certificate or TLS handshake failed. Not a site change.
No response (timeout)The server accepted nothing before the timeout. Usually hosting load, not a site change.
Blocked by firewall or CDNA firewall or CDN blocked the request. Not a site change.
Redirect loopThe site redirected endlessly. Check the site address, HTTPS and any redirect plugin — the probe never reached a page, so no specific change is implicated.
UnreachableThe failure mode could not be classified, so no cause is claimed.

Cause gating

Look at that table again. Most of those failures have nothing to do with your plugins.

So ChangeTrace refuses to rank causes for them. Only WordPress critical error and stuck in maintenance mode — the two cases where WordPress demonstrably ran and failed — get ranked causes. Every other failure mode shows "No site change is implicated" and says why.

This is deliberate, and worth defending

Suppose your host's network dies at 14:00. You happened to update a plugin at 13:50. A naive tool would rank that plugin as the likely cause, you would roll it back, and your site would still be down — because the plugin was never the problem.

ChangeTrace would rather say nothing than name the wrong thing. When you see "No site change is implicated", that is a finding: take it to your host, not your plugin list.

Some deliberate non-behaviours

  • A 404 is not an outage. Your site responded. Only 401, 403 and 429 are treated as a block.
  • Failures on our side are never recorded. If ChangeTrace cannot safely probe a URL, it logs it and records nothing, rather than reporting a site nobody could check as down.
  • Uptime status is separate from heartbeat status. They can disagree, and when they do, the outside-in check wins.

What you can configure

Nothing, currently. Monitoring is on for every active site and probes https://your-domain. If you need a different URL checked — a status endpoint, or a site behind a splash page — use the Need help? widget.

{ }For developers

Defaults: 5-minute sweep, 10s timeout, 2 failures to open, 1 success to close, 30 days of check retention. Only wp_fatal and maintenance_stuck are cause-attributable; the correlation job skips everything else both in its sweep query and per-incident. Targets resolving to private, loopback or link-local addresses are refused as an SSRF guard, which is also why probing a .test domain locally needs an explicit opt-in.

On this page