Blog / Engineering

Why We Moved Mythex's Servers to ARM

Once heavy work moved into per-project sandboxes, our servers used a third of their memory. Smaller ARM instances cut costs by about 30%. Here's how.

Mythex Team · 2026-09-29 · 2 min read

On September 19 we moved Mythex's production backend from x86 servers to smaller ARM-based AWS Graviton servers, and moved the platform database to a Graviton instance too. The trigger was simple: after we moved the heavy work — running and building each user's app — into isolated per-project sandboxes, our main servers were peaking at about a third of their memory. Smaller ARM machines fit the real load and cut the monthly server bill by roughly 30%.

What our servers actually do now

Earlier on, the backend servers did a lot of work directly. Around mid-September we finished moving the expensive parts — running each project's code, installing packages, building apps — into private cloud sandboxes, one per project. That's also what gives every project its own isolated environment and live preview.

What stayed on our own servers is the coordination: the API, the orchestrator that manages sandboxes and publishing, authentication, billing, and the gateway. That work is steady and light. After the move, the servers were peaking around 1.3 GB of memory on machines with about three times that.

Why ARM

AWS's Graviton processors are ARM chips designed by AWS. For the same money you generally get more capacity than a comparable x86 instance, or the same capacity for less — and for a workload like ours (Node.js services, Docker, Postgres) there's nothing tying us to x86.

We moved:

  • Two app servers from mid-size x86 instances to small Graviton instances with 2 GB of RAM each.
  • The platform database from a small x86 database instance to its Graviton equivalent.

The one real complication: images for two architectures

Our staging environment still runs on a small x86 server. Production now runs ARM. A container image built for one won't run on the other.

So every image is now built for both architectures under one tag — a multi-arch image. Docker picks the right variant on each machine automatically, and the deploy pipeline doesn't need to know which environment it's pushing to. The cost is a longer CI build; the benefit is that staging and production run exactly the same code from exactly the same tag.

Planning the rollback before the move

Changing CPU architecture under a live service is the kind of change you want to be able to undo in minutes. Before switching, we:

  • kept the previous server template version as a named rollback target,
  • took a database snapshot immediately before resizing the database, and
  • wrote a runbook covering the move and the way back.

We didn't need the rollback. We'd still do it the same way.

What we'd watch for

Smaller servers have less headroom. Two things follow from that:

  1. If heavy work ever moves back onto these servers, 2 GB won't be enough. The fix then is a larger Graviton size, not going back to x86.
  2. Small machines need swap. A week later one of these servers froze from memory pressure without technically running out of memory — the story of the missing swap file.

What we took from it

  • Re-measure after architecture changes. Moving work into sandboxes changed what our servers needed; the instance sizes hadn't caught up.
  • ARM is a low-risk win for typical web stacks once your images are multi-arch.
  • Write the rollback first. It makes the change calmer even if you never use it.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.

Start building free · Templates · Docs