Pressing Publish Twice Now Publishes Once
Three publish requests for one app arrived 84ms apart and each ran a full deploy. How we made the database decide who publishes, and told the others the truth.
Mythex Team · · 4 min read
Three publish requests for the same app reached our API 84 milliseconds apart, and each one started a full deploy of that app. They raced each other onto the same hosting, several of the resulting releases failed because another had got there first, and the only visible sign was a "release failed" email — the app itself ended up fine. Now a publish that arrives while another is running for the same app doesn't start a second build. The database decides which request wins, and the others are told the truth: a publish is already running.
What happened
When you publish an app on Mythex, the API records a deployment for it and asks our orchestrator — the service that runs builds — to build and release it. Publishing can start from the Publish button, from another browser tab, or from the agent when you ask it to publish. A double click, two tabs, or an agent request that arrives twice can all produce more than one publish for the same app at nearly the same moment.
In this case, three arrived within 84 ms. Each ran a full deploy. On the hosting platform, releases for the same app were created in the same second, and three of them failed because the others were already in progress. A later release stuck, and the app worked afterwards.
That's what made it easy to miss. The user saw one working app. The failures only showed up as release-failure emails.
Why nothing stopped it
The publish code did have a step that looked at in-progress deploys for the project. But what it did was mark every in-progress deploy as failed before starting — "superseded by a new publish". It was meant to clean up builds left behind by a crashed process.
In practice, it meant a second publish arriving mid-build declared the first one dead and started its own. Both ran to completion against the same app. There was no lock; there was a step that actively cleared the way for a race.
There was a second, stranger symptom for an app's very first publish. With no deployment record to reuse yet, two concurrent first publishes both tried to insert one. The database has a unique index on the app's link name (its slug, like your-app in your-app.mythex.ai), so the second insert failed. The API translated that into "that link name is taken" — advising the user to rename an app that was, at that moment, publishing successfully under that name.
The fix: the claim is the lock
Claiming the deployment record is now a single conditional database update: take this row and mark it building only if nobody is publishing it. Because it's one statement, the database settles requests that arrive in the same millisecond. Exactly one of them changes the row. The others see that zero rows changed, stop, and never reach the orchestrator.
A lock also needs a way out when its holder dies. If the API restarted in the middle of a publish, the row would stay marked "building" forever and the app would be unpublishable. So a row whose owner has stopped sending heartbeats — the running publish updates the row every 30 seconds, and five minutes of silence counts as dead — can be claimed again. That threshold comes from the heartbeat, not from a guess at how long a build takes. Building a user's own Dockerfile can legitimately run for a long time, and a slow build that's still heartbeating keeps its claim.
The step that failed every in-progress deploy no longer runs before a publish. Only rows a dead process left behind are cleared.
Telling the loser the truth
The request that loses now gets a clear answer: publish_in_progress, along with the ID of the deployment that's running. Each place that can start a publish handles it:
- The workspace follows the build that's already running instead of showing an error over it.
- The agent is told a publish of this app is already running and to wait for it, rather than being told something failed — which previously could send it off to "fix" a Dockerfile and try again.
- A first-publish collision on the link name is checked: if the row holding the name is this project's own running publish, the answer is "already publishing", not "that name is taken". A real collision with another project still gets the rename message.
How we checked it
We wrote a reproduction script that runs three publishes of one app at the same time against a real database and a stub orchestrator, and counts how many deploys the orchestrator is asked for. It covers a first publish and a republish of a live app. Before the fix, it counted two deploys. After, one.
The script also covers the escape hatch: it marks the app's row as building with no heartbeat for ten minutes, publishes again, and checks that the new publish claims the row and runs.
Unit tests hold the lock to the two things a lock must do: of three simultaneous republishes, exactly one gets through and the other two are told a publish is running; and a build whose owner died can be claimed. Another test checks that a new attempt doesn't carry over the previous attempt's error message.
What we took from it
- Cleanup code can be a race in disguise. "Fail anything in progress, then start" reads like hygiene. With two callers, it's each one cancelling the other.
- Let the database arbitrate. A read followed by a write leaves a gap for another request. A single conditional update doesn't.
- Every lock needs an expiry tied to liveness. Without one, a crash turns into a stuck app.
- Error messages are instructions. "That name is taken" told people to do the wrong thing. The right message was simpler and true.
More on publishing in the docs: how to publish. And a related fix: your published app stays up while you republish it.