Moving to Bigger Workspaces Without Making Anyone Wait
A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
Mythex Team · · 4 min read
Every Mythex project gets its own workspace: a small cloud computer where the agent writes code and your app's preview runs. We had raised new workspaces from 1 GB of memory to 2 GB, but older ones stayed at 1 GB — on the day we checked, about a quarter of them. Our first fix moved a project to a bigger workspace when it was next opened. On its first day in production, that made one person wait seven minutes. We put a hard time limit on it, then took the move out of opening a project altogether.
Why old workspaces stayed small
Our sandbox provider sets a workspace's size when it is created, and keeps it through every pause and resume. Changing the template — the recipe new workspaces are built from — only affects workspaces created after the change. Nothing resizes the ones that already exist.
That mattered. One of the small workspaces ran out of memory during a git push while its preview was starting, and stopped responding for 11 minutes. A publish failed because of it.
First attempt: move the project when it wakes
Workspaces pause when nobody is using them and wake when the project is opened. Waking seemed like the right moment to move a project, because nobody is working in it yet.
The move worked like this:
- Check the size. If the workspace is smaller than 2 GB, move it. If its size is unknown, do nothing.
- Don't disturb a publish. A publish reads the workspace's files, so if one is running, skip the move.
- Pack the project. Archive everything in the project — including its git history, which is every checkpoint the project has, and the app's own
.envfile, which belongs to the owner — but leave out installed dependencies (node_modules), which can be reinstalled. - Copy it straight across. Stream the archive from the old workspace into a fresh 2 GB one, so it never has to fit in our own server's memory.
- Compare before switching. Count the files, read the latest checkpoint, and count the checkpoints, in both workspaces. Switch only if all three match.
- Keep the old one on any failure. Any error, or any mismatch, throws away the new workspace and opens the project in the old one, as before.
One detail in the comparison mattered more than it looks. The check runs as the administrator account, and git refuses to read a repository owned by a different user. It printed nothing on both sides — and "nothing" matched "nothing", so the check passed without comparing any history at all. We told git to trust the directory, and added a test so that can't slip back.
What went wrong on day one
On its first day in production, packing one project hit a five-minute command timeout — inside the wake. The owner waited seven minutes for their workspace to open. And because the move had failed, the project was still small, so the next open would have tried again and made them wait again.
A migration that makes someone wait while opening their own project is a bad trade. The whole point of doing it on wake was that nobody would notice.
Second attempt: a hard budget
The same day we changed four things:
- One minute for the whole move. Every step — checking the size, packing, creating the new workspace, copying, unpacking, comparing — shares a single 60-second budget. When it runs out, the project opens in the old workspace.
- Stop the work, not just the waiting. Giving up on our side isn't enough if the archiving is still running inside the workspace the owner is about to use. Each archive command is now killed inside the workspace itself shortly before the budget ends. If the new workspace arrives after we've given up, it is thrown away.
- Don't make the owner wait for cleanup. Deleting the temporary archive now happens in the background.
- Don't try again on every open. A project that couldn't be moved is left alone for 24 hours, so a failure costs at most one delay a day, not one per open.
Third attempt: take it out of the wake
That made the worst case one minute. But a minute is still a minute, and we realised the move didn't need to be on anyone's path at all.
No workspace is created at 1 GB any more. So this isn't an ongoing process; it's a one-time job with a fixed list of workspaces. Doing it inside the wake meant a real person paid for every slow move, and the only way we'd learn a workspace wouldn't move was that someone had just waited for it.
So we removed the move from waking entirely. Opening a project does what it did before. The remaining small workspaces are handled as a supervised batch instead: nobody is waiting on it, and a workspace that won't move shows up in the batch's report, where we can look at it, rather than in front of the person who owns it.
What we took from it
- Resizing isn't always possible. When a platform fixes a size at creation, raising the default only helps new things. Plan for the old ones.
- "Nobody's using it yet" still means somebody is waiting. Work done on the way into a product is work a person watches.
- Put a budget on anything in a user's path, and stop the work when it runs out. Timing out on our side while a command keeps running elsewhere only moves the problem.
- A one-time migration belongs in a one-time job. Hiding it inside everyday actions makes it cheaper to write and more expensive to run.
- Test your safety checks. A comparison that compares nothing looks exactly like one that passed.
If you build on Mythex, there's nothing to do: the aim is simply fewer workspaces running out of memory, with opening a project as fast as before. How workspaces pause and wake is covered in the docs under private workspace. For another out-of-memory story, see the out-of-memory crash hiding in a progress label.