galena

Quorum, hysteresis and an approval gate

Why detection waits for agreement, and why monitor-drafted incidents will wait for a person.

Problem

A false public incident damages trust more than one that is slightly late. Single-region blips, network trouble on the probe's side, and flapping endpoints are all common.

Decision

  • A region is failing after 2 failed checks in a row, and healthy again after 3 passing ones.
  • A monitor is down when at least 2 trustworthy regions are failing. Regions whose canary fails, or that go silent, don't vote.
  • More than 4 state changes in 30 minutes marks a monitor flapping, which holds it still until it settles.
  • Incidents drafted from monitors will follow each monitor's publish policy. The default, "ask a person first", waits 10 minutes for someone to approve, then publishes if the monitor is still down. Because the name suggests nothing goes out without a person, the policy picker and every approval say that deadline in words.

All of these are per-monitor settings. See Detection.

Detection is built as described. Drafting incidents from monitors, and the approval that goes with it, are not built yet: today a monitor changes its component's status, and a person posts the incident.

Alternatives

OptionWhy not
Any region failingNoisy
Every region failingMisses regional outages
Publish everything automaticallyFalse positives go public
Never publish automaticallySlow at night; the approval timeout fails open instead

Consequences

  • With 60-second checks, the fastest detection is about two minutes.
  • The replay harness, which runs recorded results through the detection code, is the regression test for every change to these rules.

Revisit when

Replay data shows missed or late detections, or people need detection faster than a minute.

On this page