Blog / Engineering

Catching Apps That Crash Right After Publishing

Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.

Mythex Team · 2026-09-29 · 4 min read

On September 25, three publishes were marked successful even though the app's API crashed seconds into every boot. Because they "succeeded", no rollback ran; a background check failed them a few minutes later, and the whole site went to a 404. The health check was counting crashes in the hosting platform's event list — which only ever holds the last five events per machine, so a crash loop keeps it full and the count never grows. It now compares the time of the newest crash instead, and a crash-looping publish fails and restores the previous version.

What a publish waits for

When you publish an app with a backend on Mythex, the new version is rolled onto machines on our hosting platform, and a health check decides whether the publish went live or failed. If it failed, a rollback puts the previous version back.

The hard part is that "started" isn't the same as "serving". The platform reports a machine as started as soon as it's up. An app that crashes a second into boot is restarted immediately, and a moment later it reads as started again. So the health check had grown a few layers over time:

  1. Look twice. After a machine reads as started, wait briefly and look again. A crash on boot often fails the second look.
  2. Watch for crash loops. If the platform reports that a machine kept crashing and it gave up restarting it, fail.
  3. Count crashes since the publish began. Two "started" looks can straddle a restart. So if any crash happened since this publish began, look a third time; if there are more crashes than on the previous look, the app isn't staying up.

The third check is the one that failed here.

Why the count stayed flat

The crash information comes from each machine's event list: starts, exits, restarts. The platform keeps only the last five events per machine.

Picture an app that crashes 2.5 seconds into every boot, as one of them did. Every cycle adds a crash and a start. After a couple of cycles, the list is full. From then on, each new event pushes an old one out. The list slides forward, but it always holds roughly the same mix: a couple of crashes, a couple of starts.

So "how many crashes since the publish began" returned the same number on every look. The second look saw, say, two crashes. The third look also saw two — different crashes, but the same count. The check concluded the crashing had stopped, and passed the publish.

That happened three times on September 25, across two apps. Each time:

  • the publish was marked live,
  • no rollback ran, because nothing had failed,
  • our reconciler, a background job that checks deployments after the fact, failed the deployment a few minutes later,
  • and the whole site returned 404.

A publish that fails on time is recoverable: the rollback restores the previous version. A publish that fails a few minutes late, after being declared a success, has already skipped its rollback.

How we found it

We reproduced it locally: took a working API, let it go to sleep, then republished it with a change that crashed on boot. The health check marked it live.

That matched production: in each of the three publishes, the app died seconds into every boot, yet the crash count the check compared never went up between looks.

The fix: compare the newest crash, not the count

Instead of counting crashes, the check now looks at the time of the most recent crash since the publish began.

If the second look shows any crash since the publish started, the check waits and looks a third time. If the newest crash on the third look is later than the newest crash on the second, the machine crashed again in between. It's not staying up, and the publish fails with a message saying the machine crashed repeatedly since publishing and is still restarting.

This works no matter how short the event list is. A sliding window of five events can hide how many crashes there were, but it can't hide that a new one just happened — the newest event is always the one kept.

With this change, the local reproduction behaves correctly: the publish fails, and the previous image is restored. There's also a test that simulates the platform's behaviour — every look adds a crash and a start and trims the list to five — and checks the verdict is a failure, not a pass.

What we took from it

  • Know the shape of the data you're counting. A count over a capped list stops meaning anything once the list is full. The cap wasn't hidden; we just hadn't designed the check around it.
  • Prefer "did something new happen?" over "how many happened?" The newest timestamp survives truncation. A total doesn't.
  • A late failure is worse than an early one. The same broken build is harmless if the publish fails and rolls back, and takes a site down if it's declared a success first.
  • Reproduce with the real conditions. The bug needed a machine that had already crashed enough times to fill its event list. A single quick crash in a test wouldn't have shown it.

This fix pairs with another from the same day, which made sure the rollback puts the old version back rather than deleting it: a failed republish never deletes a working backend. If your own app fails to publish because it crashes on start, the publish troubleshooting docs and our guide on debugging an AI-built app are good places to start.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.
  • Keeping Long Agent Turns Inside the Context Window — Our sub-agents could grow until the model refused them, stop early to save room, or miscount tokens in Russian. Three fixes that let long AI work finish.

Start building free · Templates · Docs