Blog / Engineering

Moving to Bigger Workspaces Without Making Anyone Wait

A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.

Mythex Team · 2026-09-29 · 4 min read

Every Mythex project gets its own workspace: a small cloud computer where the agent writes code and your app's preview runs. We had raised new workspaces from 1 GB of memory to 2 GB, but older ones stayed at 1 GB — on the day we checked, about a quarter of them. Our first fix moved a project to a bigger workspace when it was next opened. On its first day in production, that made one person wait seven minutes. We put a hard time limit on it, then took the move out of opening a project altogether.

Why old workspaces stayed small

Our sandbox provider sets a workspace's size when it is created, and keeps it through every pause and resume. Changing the template — the recipe new workspaces are built from — only affects workspaces created after the change. Nothing resizes the ones that already exist.

That mattered. One of the small workspaces ran out of memory during a git push while its preview was starting, and stopped responding for 11 minutes. A publish failed because of it.

First attempt: move the project when it wakes

Workspaces pause when nobody is using them and wake when the project is opened. Waking seemed like the right moment to move a project, because nobody is working in it yet.

The move worked like this:

  1. Check the size. If the workspace is smaller than 2 GB, move it. If its size is unknown, do nothing.
  2. Don't disturb a publish. A publish reads the workspace's files, so if one is running, skip the move.
  3. Pack the project. Archive everything in the project — including its git history, which is every checkpoint the project has, and the app's own .env file, which belongs to the owner — but leave out installed dependencies (node_modules), which can be reinstalled.
  4. Copy it straight across. Stream the archive from the old workspace into a fresh 2 GB one, so it never has to fit in our own server's memory.
  5. Compare before switching. Count the files, read the latest checkpoint, and count the checkpoints, in both workspaces. Switch only if all three match.
  6. Keep the old one on any failure. Any error, or any mismatch, throws away the new workspace and opens the project in the old one, as before.

One detail in the comparison mattered more than it looks. The check runs as the administrator account, and git refuses to read a repository owned by a different user. It printed nothing on both sides — and "nothing" matched "nothing", so the check passed without comparing any history at all. We told git to trust the directory, and added a test so that can't slip back.

What went wrong on day one

On its first day in production, packing one project hit a five-minute command timeout — inside the wake. The owner waited seven minutes for their workspace to open. And because the move had failed, the project was still small, so the next open would have tried again and made them wait again.

A migration that makes someone wait while opening their own project is a bad trade. The whole point of doing it on wake was that nobody would notice.

Second attempt: a hard budget

The same day we changed four things:

  • One minute for the whole move. Every step — checking the size, packing, creating the new workspace, copying, unpacking, comparing — shares a single 60-second budget. When it runs out, the project opens in the old workspace.
  • Stop the work, not just the waiting. Giving up on our side isn't enough if the archiving is still running inside the workspace the owner is about to use. Each archive command is now killed inside the workspace itself shortly before the budget ends. If the new workspace arrives after we've given up, it is thrown away.
  • Don't make the owner wait for cleanup. Deleting the temporary archive now happens in the background.
  • Don't try again on every open. A project that couldn't be moved is left alone for 24 hours, so a failure costs at most one delay a day, not one per open.

Third attempt: take it out of the wake

That made the worst case one minute. But a minute is still a minute, and we realised the move didn't need to be on anyone's path at all.

No workspace is created at 1 GB any more. So this isn't an ongoing process; it's a one-time job with a fixed list of workspaces. Doing it inside the wake meant a real person paid for every slow move, and the only way we'd learn a workspace wouldn't move was that someone had just waited for it.

So we removed the move from waking entirely. Opening a project does what it did before. The remaining small workspaces are handled as a supervised batch instead: nobody is waiting on it, and a workspace that won't move shows up in the batch's report, where we can look at it, rather than in front of the person who owns it.

What we took from it

  • Resizing isn't always possible. When a platform fixes a size at creation, raising the default only helps new things. Plan for the old ones.
  • "Nobody's using it yet" still means somebody is waiting. Work done on the way into a product is work a person watches.
  • Put a budget on anything in a user's path, and stop the work when it runs out. Timing out on our side while a command keeps running elsewhere only moves the problem.
  • A one-time migration belongs in a one-time job. Hiding it inside everyday actions makes it cheaper to write and more expensive to run.
  • Test your safety checks. A comparison that compares nothing looks exactly like one that passed.

If you build on Mythex, there's nothing to do: the aim is simply fewer workspaces running out of memory, with opening a project as fast as before. How workspaces pause and wake is covered in the docs under private workspace. For another out-of-memory story, see the out-of-memory crash hiding in a progress label.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.
  • Keeping Long Agent Turns Inside the Context Window — Our sub-agents could grow until the model refused them, stop early to save room, or miscount tokens in Russian. Three fixes that let long AI work finish.

Start building free · Templates · Docs