galena

SQS FIFO grouped by monitor

Why check results travel through one ordered queue.

Problem

Results for one monitor arrive from several regions at nearly the same moment. Evaluating them concurrently risks lost updates and duplicate transitions. A probe may also send the same result twice.

Decision

One SQS FIFO queue in the home region. The message group is the monitor's id, so each monitor's results are evaluated one at a time, in order. The deduplication id is {monitorId}#{region}#{epochMinute}, with the minute taken from the scheduler's scheduled time, so a repeated send within five minutes is dropped. The monitor's state item also carries a version that every write is conditional on, as a second guard.

Alternatives

OptionWhy not
A standard queue and DynamoDB locksMore code and more ways to fail
KinesisShards to manage and pay for, for no benefit at this size
Probes write to DynamoDB directlyDetection would run in every region separately

Consequences

  • FIFO throughput limits apply; they are far above what a deployment needs.
  • A message that keeps failing blocks its monitor's group until it moves to the dead-letter queue, after five attempts.

Revisit when

Sustained throughput approaches FIFO limits.

On this page