← Blog
✎blog post

A Reboot Is Not a Health Check

A dark control panel showing a row of green status lights, a restart key, and a browser route map reflected on metal.

The production stack came back from a reboot with 37 of 37 checks passing only after the check covered containers, Cloudflare routes, and the browser analytics path.

The post-reboot result was 37/37. That number included every expected production container, cloudflared, fourteen public routes through Cloudflare, and a browser-backed consent to analytics check across three sites. A local process returning 200 would have missed most of what visitors depend on.

The stack is a path, not a process

The check starts with the expected container names. It then verifies cloudflared is active, because a healthy app behind a dead tunnel is still down to the public internet. Next it requests the public URLs through their real domains, including the redirect that protects the operator panel.

The last assertion runs the consent to tracker to stored-event path. Analytics is not proven by loading a JavaScript file. The check must see the browser event travel through the deployed stack and land where the analytics service expects it.

A green local health endpoint can lie by omission

Every layer can be healthy in isolation while the user path fails. The app can answer locally while Caddy points at an old slot. Cloudflared can run while a route returns the wrong status. The tracker can load while consent prevents the event, or the event can be accepted while storage is broken.

The 37 checks are intentionally repetitive. Repetition is useful when each assertion names a different boundary. The script prints the failing name and exits nonzero rather than turning a partial boot into a green deployment.

Check the path visitors use

After a reboot, prove the public path, not the process list.

What counts:

  • expected containers running;
  • tunnel active;
  • public domains returning the expected status;
  • browser consent and analytics storage succeeding;
  • a fail count that identifies the broken boundary.

What does not count:

  • docker ps alone;
  • a local health endpoint without the public host;
  • a 200 from the home page as a proxy for every route;
  • a JavaScript file returning 200 without an event receipt.

The cost was a longer reboot gate

The check takes longer than docker ps and depends on the public edge, browser tooling, and analytics state. That cost is the point. A short check that cannot touch the failure boundary buys a fast false positive.

It also creates maintenance work. The expected container list and public route list must change with the stack. That list is a useful inventory, not disposable test code.

How to prove this wrong

Run prod/ops/reboot-check.sh on the production host after a reboot. Read every OK line, not only the final count. Break one layer at a time in a controlled window: stop a container, disable a route, or block analytics storage. The script should identify the corresponding check and exit nonzero. Restore the layer, rerun, and require 37/37.

Provenance

The check is prod/ops/reboot-check.sh at commit 274fd62. It covers 21 containers, cloudflared, fourteen public route assertions, and one consent to analytics check, for 37 total. The production host used one local stack and one public edge. This proves that reboot, not every future deploy, passed this exact gate.

Discuss this on gllm.forum.

Comments

No comments yet.

Log in to comment.