Blog / Engineering

How We Got Deploys Down to Zero 502 Errors

A watchdog racing our deploys and a load balancer that noticed too late: how we found both, fixed them, and measured deploys from 10 502s to zero.

Mythex Team · 2026-09-29 · 4 min read

Every time we deployed Mythex's backend, a handful of requests failed with 502 Bad Gateway. Not many — but a 502 during a deploy means someone's click did nothing, or a build they were watching stalled for a moment. We traced it to two separate problems: a self-healing watchdog that fought our deploys, and a gap between when a server stopped listening and when the load balancer stopped sending it traffic. Across three consecutive production deploys after the fixes, the count went from 10 to 6 to 0.

The setup

Our API and orchestrator run as containers on a small group of servers behind an AWS Application Load Balancer (ALB). A deploy updates one server at a time: it drains the old containers, starts new ones, and moves on. A watchdog on each server checks that the containers are up and restarts anything that isn't — it exists so a crashed process at 3 a.m. comes back without anyone waking up.

Both of those are reasonable on their own. Together they weren't.

Problem 1: the watchdog raced the deploy

The watchdog ran once a minute and "healed" on one failed health check. A deploy deliberately drains services for up to 150 seconds — during which a health check is supposed to fail. So the watchdog fired during every single deploy.

Most of the time it lost the race harmlessly. Once, it won: it rewrote the compose file (the description of which containers should run) in the middle of a deploy, and left a server with no gateway container at all — while the deploy job reported success.

The fixes:

  • One lock. The deploy now takes the same file lock as the server's bootstrap script, so the watchdog and a deploy can never run at the same time.
  • Four strikes, not one. The watchdog needs four consecutive failed checks before it acts. A draining service no longer looks dead.
  • One definition. The compose file now has exactly one source, published to the servers, instead of copies that could drift.
  • Verify, don't trust. The deploy job no longer passes because a command returned. It checks every server's API, orchestrator, gateway and publish worker before it can succeed.

Problem 2: the server left before the traffic did

The second problem was quieter. When a deploy drains a server, the server starts answering its health endpoint with 503 so the load balancer takes it out of rotation. Ours did that for 5 seconds, then shut down.

But the ALB only checked health every so often, and needed several failed checks in a row before it stopped routing to a target — up to about 60 seconds end to end. So for roughly 55 seconds, the load balancer kept sending real requests to a server that had already stopped listening. Those were our 502s.

The fix was to make the two timelines agree:

  • The target group now checks every 5 seconds and marks a target unhealthy after 2 failures — about 10 seconds to notice.
  • The drain keeps serving real traffic for 15 seconds after it starts failing health checks, so there's always overlap.
  • These settings live in the script that provisions the infrastructure, not in a console someone clicked once.

Making the numbers hold each other

Timeouts like these depend on each other. The drain has to outlast the time the balancer needs to notice; the drain plus shutdown has to finish before the container is force-killed. It's easy for someone to change one number in six months and quietly reintroduce the bug.

So we wrote those relationships down as tests: the drain must be longer than the health-check interval times the failure threshold; the time to be marked healthy again must stay short; the whole drain must fit inside the container's stop grace period. Change one number without the others and CI fails. We mutation-tested the tests themselves — changed each number to a wrong value and checked that a test caught it.

Measuring a deploy is easy to get wrong

Our first measurement said a clean deploy was about 3% broken. It wasn't. We were counting every non-200 response, and a 503 from a health endpoint during a drain isn't a failure — it's the drain doing its job, telling the balancer not to pick that server while it keeps serving real requests.

What users feel is a 502: the request reached nothing. Counting only those gave the real picture — 10, then 6, then 0 across three production deploys.

What we took from it

  • Self-healing needs to know about planned downtime. A watchdog that can't tell a deploy from a crash will eventually fight the deploy.
  • Draining is a handshake. The server has to keep serving until the load balancer has provably stopped sending traffic, and that timing is set on the balancer, not the server.
  • Encode invariants as tests. Operational numbers drift. Tests that relate them to each other don't.
  • Decide what you're measuring before you measure. "Errors" was the wrong metric; "requests that reached nothing" was the right one.

If you publish an app on Mythex, none of this is visible to you — which is the point. Related: how your published app stays up while you republish it.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.

Start building free · Templates · Docs