galena
Architecture

Cold path

How a change in the dashboard becomes a published page and a notification, exactly once.

The cold path handles everything people do and everything that follows a state change. It is built on two rules: every change is recorded in the same transaction as the event that announces it, and everything that can run twice carries a deterministic key.

The API

The API is Hono on one Lambda behind an API Gateway HTTP API, throttled at 50 requests a second with bursts of 100. The dashboard reaches it on its own origin: the dashboard's CloudFront distribution forwards /auth/* and /v1/* to it, so cookies are first-party and there is no CORS. A secret header that only CloudFront adds keeps the API from being called directly.

Requests are validated with the same Zod schemas the dashboard uses, and errors follow RFC 9457 (application/problem+json) with a stable code. Domain rules live in packages/core; the API calls them through repositories in packages/db.

Domain writes and the outbox

Every change (an incident update, a new monitor, a reordered group) runs in one database transaction that writes:

  1. the change itself;
  2. an audit log entry;
  3. an outbox row describing the event, such as incident.updated or monitor.changed.
A change and its outbox row The API writes a change and its outbox row in one transaction, then triggers the dispatcher, which reads the row, hands the event to the publisher and the notifier, and marks it dispatched. save update one transaction trigger dispatch 201 read the row publish, notify mark dispatched Dashboard API Lambda Aurora Postgres Dispatch outbox.dispatch Consumers publish, notify

After the commit, the API triggers the outbox.dispatch task with the key outbox:{outboxId}. Triggering after the commit means the dispatcher never sees a change that could still roll back; writing the row inside the transaction means no committed change can lose its event.

The dispatcher reads the row, does what the event needs, and marks it dispatched:

EventDispatch does
monitor.changedRewrites monitors.json for the probes
maintenance.*Starts or replaces the window's lifecycle run, rewrites monitors.json, publishes, notifies
incident.*Publishes, notifies
component.changed, component_group.changedPublishes
subscriber.requestedSends the confirmation email

Dispatch runs one at a time. Each run rebuilds monitors.json from the database, so the newest committed state is always written last.

If trigger.dev is unreachable at the moment of the trigger, the request still succeeds and the row stays pending, but nothing dispatches it later yet. An hourly sweep of pending rows is planned.

Why nothing polls the database

Aurora Serverless v2 pauses after five idle minutes and costs nothing for compute while paused. A dispatcher that polled the outbox every minute would keep it awake around the clock. So the cold path is event-driven: the API triggers dispatch directly, and the only scheduled job that reads Aurora is the hourly uptime rollup, at five past the hour.

The first query after a pause can take around 15 seconds, and the Data API fails fast with DatabaseResumingException meanwhile. The database client retries that error for up to 30 seconds, and the dashboard explains the wait after 10.

Tasks

TaskStarted byRunsIdempotency key
monitor.state-changedThe evaluator, on a transitionRecords the transition, republishes unless suppressedmon:{monitorId}:{transitionSeq}
outbox.dispatchThe API, after each commitOne at a timeoutbox:{outboxId}
page.publishDispatch, state changes, the rollupOne at a time; latest winspub:{snapshotVersion}
page.rebuild-htmlEach publishOne at a time; latest winshtml:{snapshotVersion}
notify.fanoutDispatch, for incident and maintenance eventsfan:{eventId}
notify.emailFan-out, and subscription requestsUp to 5 attemptssend:{eventId}:{subscriberId}
notify.slack, notify.webhookFan-outUp to 5 attempts, 10 at a time per channelsend:{eventId}:{endpointId}
maintenance.lifecycleDispatch, when a window is scheduled or editedSleeps until start and endmnt:{maintenanceId}:{version}
rollup.uptimeCron, five past every hourReplays yesterday and today, then publishes

trigger.dev tasks run on trigger.dev's machines, not in AWS. They reach AWS only as the galena-<stage>-worker-access IAM user, whose policy allows exactly what the tasks need.

Notification fan-out

Notification fan-out The dispatcher hands an incident or maintenance event to notify.fanout, which writes one delivery row per subscriber or endpoint and starts one send each, by email through SES, to Slack, or to a webhook URL. event one row each one send each Dispatch outbox.dispatch Fan-out notify.fanout Deliveries Aurora Email notify.email Slack notify.slack Webhook notify.webhook SES confirm, notify Slack channel incoming webhook Your endpoint signed POST

notify.fanout decides who hears about an event (see Notifications), builds one notice, creates one delivery row per target, and triggers one send per row. Each row is unique on (event, target), so a re-run of the fan-out finds the rows already there and sends nothing twice. A send checks that its target still qualifies (still subscribed, endpoint not disabled) before sending, and records the result on its row.

On this page