galena
Concepts

Monitors and checks

What a monitor checks, what one check records, and the settings that shape both.

A monitor is an automated check of one URL. Every minute, a probe in each probe region runs it once. A monitor may be attached to a component; detection then drives that component's status. A monitor without a component is still checked and shown in the dashboard, but never changes the page.

Galena has HTTP monitors today. TCP, DNS, certificate expiry and heartbeat monitors are on the roadmap.

What an HTTP check does

  1. Sends a GET (or HEAD) to the URL from a fresh connection, as the user agent Galena-Probe (+https://github.com/astrlme/galena).
  2. Follows up to 5 redirects, unless redirects are turned off.
  3. Passes when the final status is one you accept (any 2xx by default) and, when you set a keyword, the first megabyte of the body contains it.
  4. Gives up after the timeout, 10 seconds by default.

Every address is checked by the SSRF guard twice: before the request, and again for each address DNS returns as the socket connects. Private, loopback, link-local, carrier-grade NAT, documentation and other special-purpose ranges are refused, in IPv4 and IPv6, on every redirect hop. A check aimed at one fails with blocked_by_guard. See Security model.

Settings

SettingDefaultIn the dashboard
NameYes
URL (http:// or https://, no credentials in it)Yes
Method: GET or HEADGETYes
Keyword the body must contain (GET only)noneYes
ComponentnoneYes
Down status: what the component shows while the monitor is downMajor outageYes
Publish policyAsk a person firstYes
PausednoYes
Accepted status codes, up to 20any 2xxAPI
Timeout, 1 to 10 seconds10 sAPI
Follow redirectsyesAPI
Detection settingssee DetectionAPI

Settings marked API are set with PUT /v1/monitors/{id}; see the API reference.

Publish policy

Each monitor carries a publish policy for incidents it will draft on its own:

PolicyMeaning
Ask a person first (approve)The draft waits 10 minutes for a person, then publishes if the monitor is still down.
Publish automatically (auto)The draft publishes as soon as the monitor is down.
Internal only (internal_only)The draft stays in the dashboard. Nothing reaches the page or subscribers.

Galena stores the policy today, but drafting incidents from monitors is not built yet: a monitor going down changes its component's status and republishes the page, and a person posts the incident. See the roadmap.

Check results

One probe's outcome for one monitor in one minute is a check result:

FieldMeaning
regionThe probe region that ran it
scheduledAtThe minute the scheduler fired; results are de-duplicated on it
statusup, down or error (the probe failed, not the target)
httpStatusThe final response's status code, if one arrived
latencyMsTotal time
phasesDNS, connect, TLS (null for http://), time to first byte, and total, in ms
errorWhy it failed: a code and a message
Error codeMeaning
timeoutNo complete answer within the timeout
dns_failedThe name didn't resolve
connection_failedThe connection was refused or reset
tls_failedThe TLS handshake or certificate failed
http_statusA status code you don't accept, or more than 5 redirects
keyword_missingThe body didn't contain the keyword
blocked_by_guardThe SSRF guard refused an address
probe_failedThe probe itself broke

The dashboard shows each monitor's last hour per region. Results stay in DynamoDB for 90 days. The page's history is built from something smaller: each monitor's confirmed state changes, kept in the database. See Detection.

On this page