Blog / Engineering

A Failed Republish Never Deletes a Working Backend

A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.

Mythex Team · 2026-09-29 · 4 min read

On September 25, a failed republish deleted an app's backend that had been live and working for about two and a half hours. The code that undoes a failed publish removed every server app the publish had touched — including ones that were already serving before it started. Now only apps a publish created are ever deleted; an app that already existed is put back on the exact version it ran before. Two follow-up fixes the same day made sure that undo also runs when the hosting platform's own deploy tool fails, and that a healthy app isn't failed just because it was asleep.

How a backend is published

When your Mythex app has a backend — an API or server, rather than just static files — publishing builds it from a Dockerfile (a recipe for packaging the app into an image) and rolls that image onto machines on our hosting platform. Then a health check waits for the new version to actually come up and stay up before the publish is marked live.

If anything fails along the way, a rollback runs to undo it. Getting that rollback right matters more than almost anything else in publishing, because it runs precisely when something is already going wrong.

Problem 1: rollback deleted what it should have restored

The rollback kept a list of the server apps a publish had deployed to, and on failure, destroyed every one. That's correct for a first publish: the app was created a moment ago, nothing depends on it, and deleting it leaves things as they were.

It's exactly wrong for a republish. The app already existed and was serving real visitors. On September 25, an app's API that had been live for about two and a half hours was deleted by a failed republish.

There was a second failure mode hiding in the same code. Where the delete didn't go through, the broken new image stayed on the machines — while we told the owner their current site was still live. By the time the health check can fail, the new image has already been rolled onto every machine. Doing nothing isn't the same as undoing.

The fix: destroy only what you created, restore the rest

Each publish now records, for every server app it touches:

  • whether this publish created it, and
  • if not, which image it was running just before.

On failure, each app gets one of three treatments:

  1. Created by this publish: destroy it, as before.
  2. Already existed, previous image known: redeploy the previous image with the same configuration. No build — only the image moves back.
  3. Already existed, previous image unknown: leave it alone. An existing app is never deleted, even when we can't restore it.

These rules are a small pure function with its own tests, so the decision can't drift back to "destroy everything" without a test failing.

Problem 2: the deploy tool could fail before we'd written down how to undo

Soon after that fix, we found a gap in it. The hosting platform's deploy tool runs its own smoke checks: if the new machines crash, it exits with an error. But it does that after the new image is already on the machines.

Our code only registered an app for rollback once the tool returned successfully. When the tool itself failed the deploy, the app wasn't on the rollback list yet, so nothing was restored. On our staging environment, a broken republish left the broken image serving.

The fix: register the undo before starting

The app is now registered for rollback — with whether it was created and what image it ran — as soon as it's prepared and before anything is rolled onto it. Whatever fails afterwards, the rollback knows what to put back. We verified it locally: a broken republish failed in the deploy tool's own smoke checks, and the previous version kept answering.

Problem 3: a sleeping app was failed as "never started"

Published apps on Mythex sleep when idle to save your credits. A republish leaves a sleeping machine stopped, so the health check has to wake it. It asked the platform to start the machine once, and ignored the answer.

On September 25, that one start request went nowhere. A healthy API sat stopped for the full five-minute health check and was failed with "1 machine never started (last state: stopped)".

The platform does sometimes refuse a start. While reproducing this, it returned a "lease currently held" conflict without us provoking it, and it refuses starts while a machine is being updated.

The fix: keep asking, and knock on the door

The health check now:

  • asks again to start a stopped machine every 20 seconds while it waits,
  • logs what the platform answered when it refuses, and
  • sends one ordinary request to the app, the way a visitor would — the platform's proxy wakes a sleeping machine when a request arrives.

We reproduced it on the real platform with a throwaway app whose first start was refused. Before the fix, the publish failed after 91 seconds. After, it published in 20 seconds. An app that genuinely crashes on boot still fails, as it should.

What we took from it

  • Rollback has to know the difference between new and existing. "Undo" for something you created is delete. "Undo" for something that was already there is restore. Treating them the same is how a safety mechanism deletes a working app.
  • Write down how to undo before you act. If the undo is only recorded after a step succeeds, a step that fails halfway leaves nothing to undo with.
  • Never swallow the answer to a request you depend on. One ignored "no" turned a healthy app into a failed publish.

A related fix from the same day: catching apps that crash right after publishing. For what happens when a publish fails on your side, see publish troubleshooting in the docs.

Keep reading

  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.
  • Keeping Long Agent Turns Inside the Context Window — Our sub-agents could grow until the model refused them, stop early to save room, or miscount tokens in Russian. Three fixes that let long AI work finish.

Start building free · Templates · Docs