Tag: alerting

  • Ninety-seven percent of the alerts were one monitor

    A monitor checking that certificate issuance still worked had been alternating between down and up on every single cycle, for at least four days. Roughly 540 state changes a day, each one sending a notification.

    Read on →

  • The alert fired six times and the test said it did not

    A new alert had just been added to catch a server simply dying. Proving it meant inducing the failure: stop the thing that reports the machine is alive, wait, and confirm the alert fires and the notification arrives.

    Read on →

  • The notification that killed the script

    Two scripts sit on the lab’s gateways waiting for the day the primary one dies. One moves the public DNS records to the standby. The other moves them back. Both were, until this week, incapable of telling anyone they had…

    Read on →

  • Status: alerting that can actually reach you

    Two days on the public edge and no new hardware. The alerting path is now proven end to end rather than merely deployed, which is a distinction this project keeps having to relearn. It produced four more errata along the…

    Read on →