The production stack came back from a reboot with 37 of 37 checks passing only after the check covered containers, Cloudflare routes, and the browser analytics path.
The post-reboot result was 37/37. That number included every expected production container, cloudflared, fourteen public routes through Cloudflare, and a browser-backed consent to analytics check across three sites. A local process returning 200 would have missed most of what visitors depend on.
The check starts with the expected container names. It then verifies cloudflared is active, because a healthy app behind a dead tunnel is still down to the public internet. Next it requests the public URLs through their real domains, including the redirect that protects the operator panel.
The last assertion runs the consent to tracker to stored-event path. Analytics is not proven by loading a JavaScript file. The check must see the browser event travel through the deployed stack and land where the analytics service expects it.
Every layer can be healthy in isolation while the user path fails. The app can answer locally while Caddy points at an old slot. Cloudflared can run while a route returns the wrong status. The tracker can load while consent prevents the event, or the event can be accepted while storage is broken.
The 37 checks are intentionally repetitive. Repetition is useful when each assertion names a different boundary. The script prints the failing name and exits nonzero rather than turning a partial boot into a green deployment.
After a reboot, prove the public path, not the process list.
What counts:
What does not count:
docker ps alone;The check takes longer than docker ps and depends on the public edge, browser tooling, and analytics state. That cost is the point. A short check that cannot touch the failure boundary buys a fast false positive.
It also creates maintenance work. The expected container list and public route list must change with the stack. That list is a useful inventory, not disposable test code.
Run prod/ops/reboot-check.sh on the production host after a reboot. Read every OK line, not only the final count. Break one layer at a time in a controlled window: stop a container, disable a route, or block analytics storage. The script should identify the corresponding check and exit nonzero. Restore the layer, rerun, and require 37/37.
The check is prod/ops/reboot-check.sh at commit 274fd62. It covers 21 containers, cloudflared, fourteen public route assertions, and one consent to analytics check, for 37 total. The production host used one local stack and one public edge. This proves that reboot, not every future deploy, passed this exact gate.
Discuss this on gllm.forum.
Comments
No comments yet.
Log in to comment.