How We Got Deploys Down to Zero 502 Errors
A watchdog racing our deploys and a load balancer that noticed too late: how we found both, fixed them, and measured deploys from 10 502s to zero.
Mythex Team · · 4 min read
Every time we deployed Mythex's backend, a handful of requests failed with 502 Bad Gateway. Not many — but a 502 during a deploy means someone's click did nothing, or a build they were watching stalled for a moment. We traced it to two separate problems: a self-healing watchdog that fought our deploys, and a gap between when a server stopped listening and when the load balancer stopped sending it traffic. Across three consecutive production deploys after the fixes, the count went from 10 to 6 to 0.
The setup
Our API and orchestrator run as containers on a small group of servers behind an AWS Application Load Balancer (ALB). A deploy updates one server at a time: it drains the old containers, starts new ones, and moves on. A watchdog on each server checks that the containers are up and restarts anything that isn't — it exists so a crashed process at 3 a.m. comes back without anyone waking up.
Both of those are reasonable on their own. Together they weren't.
Problem 1: the watchdog raced the deploy
The watchdog ran once a minute and "healed" on one failed health check. A deploy deliberately drains services for up to 150 seconds — during which a health check is supposed to fail. So the watchdog fired during every single deploy.
Most of the time it lost the race harmlessly. Once, it won: it rewrote the compose file (the description of which containers should run) in the middle of a deploy, and left a server with no gateway container at all — while the deploy job reported success.
The fixes:
- One lock. The deploy now takes the same file lock as the server's bootstrap script, so the watchdog and a deploy can never run at the same time.
- Four strikes, not one. The watchdog needs four consecutive failed checks before it acts. A draining service no longer looks dead.
- One definition. The compose file now has exactly one source, published to the servers, instead of copies that could drift.
- Verify, don't trust. The deploy job no longer passes because a command returned. It checks every server's API, orchestrator, gateway and publish worker before it can succeed.
Problem 2: the server left before the traffic did
The second problem was quieter. When a deploy drains a server, the server starts answering its health endpoint with 503 so the load balancer takes it out of rotation. Ours did that for 5 seconds, then shut down.
But the ALB only checked health every so often, and needed several failed checks in a row before it stopped routing to a target — up to about 60 seconds end to end. So for roughly 55 seconds, the load balancer kept sending real requests to a server that had already stopped listening. Those were our 502s.
The fix was to make the two timelines agree:
- The target group now checks every 5 seconds and marks a target unhealthy after 2 failures — about 10 seconds to notice.
- The drain keeps serving real traffic for 15 seconds after it starts failing health checks, so there's always overlap.
- These settings live in the script that provisions the infrastructure, not in a console someone clicked once.
Making the numbers hold each other
Timeouts like these depend on each other. The drain has to outlast the time the balancer needs to notice; the drain plus shutdown has to finish before the container is force-killed. It's easy for someone to change one number in six months and quietly reintroduce the bug.
So we wrote those relationships down as tests: the drain must be longer than the health-check interval times the failure threshold; the time to be marked healthy again must stay short; the whole drain must fit inside the container's stop grace period. Change one number without the others and CI fails. We mutation-tested the tests themselves — changed each number to a wrong value and checked that a test caught it.
Measuring a deploy is easy to get wrong
Our first measurement said a clean deploy was about 3% broken. It wasn't. We were counting every non-200 response, and a 503 from a health endpoint during a drain isn't a failure — it's the drain doing its job, telling the balancer not to pick that server while it keeps serving real requests.
What users feel is a 502: the request reached nothing. Counting only those gave the real picture — 10, then 6, then 0 across three production deploys.
What we took from it
- Self-healing needs to know about planned downtime. A watchdog that can't tell a deploy from a crash will eventually fight the deploy.
- Draining is a handshake. The server has to keep serving until the load balancer has provably stopped sending traffic, and that timing is set on the balancer, not the server.
- Encode invariants as tests. Operational numbers drift. Tests that relate them to each other don't.
- Decide what you're measuring before you measure. "Errors" was the wrong metric; "requests that reached nothing" was the right one.
If you publish an app on Mythex, none of this is visible to you — which is the point. Related: how your published app stays up while you republish it.