Alerts That Fire Once, Not Once per Server
A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
Mythex Team · · 4 min read
One of our operational alerts — "low credits" for a builder's published app — is meant to reach us at most once a day. One afternoon it arrived five times. The cause: each copy of our API remembered which alerts it had already sent in its own memory, so two servers sent it twice, and every deploy wiped that memory and let them send it again. The same pattern had a background job doing its work twice. Both now coordinate through a shared store, and both fail on the safe side if that store is unavailable.
The symptom
Mythex sends its team alerts when something needs a look: a service is down, a job stalled, a builder's credit balance is running low while their app is published. Each alert has a dedupe window — a period during which the same alert is sent only once. For the low-credits alert, that window is 24 hours.
Five copies of the same alert in one afternoon meant the window wasn't holding. Worse, they came in pairs, roughly 11 minutes apart.
Why it happened
Our alerts package kept its dedupe in memory: a table of "alert key → when it was last sent", inside the running process. That works when there is one process that never restarts. Production has neither:
- Two API replicas. Production runs two copies of the API. Each had its own table, so each sent the alert once. That's the pairs.
- Deploys. Every deploy replaces both replicas with fresh processes and empty tables. After a deploy, the alert went out again, as if it had never been sent.
The ~11-minute gap between the two alerts in each pair came from the job that raises this alert. It runs on both replicas and takes a lock, so the replicas win it in turn, and each one raised the alert when its turn came.
There was a second, quieter bug in the same code. To stop the table growing forever, it pruned itself once it held 500 keys — deleting every entry older than 15 minutes, whatever its window. Under load, a 24-hour dedupe could last a quarter of an hour.
The fix
One shared record per alert
Before the API sends an alert, it now claims the alert's dedupe key in Redis — a shared in-memory store that both replicas talk to and that doesn't restart when the API does. The claim is a single "set this key only if it doesn't exist, and expire it after the window" operation. Whichever replica claims it first sends the alert; the other finds the key taken and stays quiet. A new process after a deploy finds the key still there.
The key and window come from the same rule the alerts package uses, now shared between them, so the two layers can't disagree about what counts as "the same alert".
Fail open, not closed
The most important decision was what to do when Redis can't be reached. We chose to send the alert anyway. If Redis is down, or doesn't answer within 1.5 seconds, the alert goes out. A duplicate alert is noise. A dropped alert is someone not being told.
The most serious alerts, P0, are never held back by dedupe at all.
Each entry keeps its own window
In the alerts package's own in-memory dedupe, each entry now stores when its window ends, and pruning removes only entries whose window is over. A 24-hour key stays for 24 hours, however many other alerts arrive.
The same bug, somewhere else
The same pattern had doubled a piece of background work too: a reconciliation job that keeps our record of each workspace's status in line with our sandbox provider's. Each pass lists every sandbox in the account — thousands in production, about five seconds — plus every running or paused project. Both replicas ran it every two minutes.
Its writes are idempotent — running them twice gives the same result as running them once — so nothing broke. It was just twice the work. Now each two-minute window is claimed in Redis the same way our usage meters already claimed theirs, and only the replica that claims it runs the pass. If Redis can't be reached, the pass runs anyway: doing the work twice is harmless, not doing it isn't.
How we tested it
The tests model the real setup. They load the alert module twice, as two separate "replicas", point both at one shared store, raise the same alert from each, then load a third fresh copy to stand in for "after a deploy" and raise it again. Exactly one alert is sent. Other tests check that a different builder's alert still goes out, that a P0 is sent from both replicas, and that the alert is still sent when the store is down or hangs past the timeout. The prune fix has its own test: a 24-hour key survives 600 other alerts arriving 20 minutes later.
For the reconciliation job, two replicas ticking at once, plus a third tick in the same window, produce one listing.
What we took from it
- In-memory state is per process. Anything that must be true across servers — "sent once", "ran once" — needs to live somewhere they share.
- Deploys are restarts. State that only survives until the next restart doesn't survive a normal working day.
- Decide which way to fail. For alerts, a shared store being down should mean more alerts, not fewer.
- Look for the pattern, not just the instance. The alert was the loud case; the doubled reconciliation job was the same bug, quietly costing work.
None of this is visible from inside Mythex. For more on how we keep the platform steady, see how we got deploys down to zero 502 errors and the missing swap file. If you're new to the idea of running more than one copy of a service, our guide on what a server is covers the basics.