Troubleshooting
What to check when something doesn't happen.
Most problems show up in one of three places: the trigger.dev dashboard (every task run, its payload and its logs), CloudWatch Logs for the Lambdas (one log group each, kept a month), and the dashboard's own error messages, which name what to do.
The first request is slow
Aurora pauses after five idle minutes. The first request afterwards waits about 15 seconds while it resumes; the dashboard says so after 10 seconds. This is expected and is what keeps an idle deployment cheap. Scripts calling the API should allow at least 20 seconds.
The page didn't update
- In trigger.dev, find the
outbox.dispatchrun for the change. If there is none, the API couldn't reach trigger.dev when you saved: the API's log saysoutbox.dispatch was not triggered. The change is saved but its event is still pending; save the change again to publish. - Check the
page.publishrun that followed.supersededmeans a newer publish already covered the change. - Check
page.rebuild-htmlfor the HTML. The data files (snapshot.json) land first, so an open page updates within 30 seconds even before the HTML does.
A monitor isn't being checked
- Is it paused? Paused monitors leave
monitors.json. - After adding or changing a monitor,
outbox.dispatchrewritesmonitors.json; the probes pick it up on their next run. - The probes' log groups (
galena-<stage>-probe-<region>) show each run; a check aimed at a private address fails withblocked_by_guard.
A monitor stays Unknown
Detection waits until at least quorum regions have a verdict, which takes a couple of checks
per region. A monitor that stays unknown longer usually has regions that only return error
(the probe failing, not the target), or a region excluded because its canary fails. The
dashboard's per-region results show which.
A monitor shows Flapping
It changed state more than four times in 30 minutes. It leaves flapping once it stays up for 30
minutes, or down for 15. If the endpoint really is unstable, that is the truth; if the checks are
too strict, give it a longer timeout or a higher failThreshold.
Emails don't arrive
- Until AWS grants SES production access, SES only sends to addresses verified in SES.
- Without
GLN_EMAIL_FROMin the workers' environment, emails are written to a local folder instead of sent. - An address that bounced or complained is suppressed for good; the dashboard shows "Stopped: the address bounced or reported spam".
- Check the
notify.emailrun and its delivery row'slast_error.
A webhook endpoint shows Failing
Its last delivery failed: the dashboard shows since when. Send test shows what the receiver answers now. A 4xx other than 429 is not retried; fix the receiver, then send a test.
The page says "Status not confirmed"
The page hasn't been republished for two hours, though the hourly rollup should republish it.
Check the rollup.uptime schedule and its runs in trigger.dev, and the workers' AWS access key.
403 not_through_cloudfront
The API only answers requests that came through CloudFront. Call it on the dashboard's address, not the API Gateway URL.