Blog / Engineering

The Missing Swap File That Froze Our Servers

A production server stopped answering without crashing — no out-of-memory kill, just page-cache thrash on a box with no swap. What happened and what we changed.

Mythex Team · 2026-09-29 · 3 min read

On September 24, one of our production servers stopped answering. The load balancer marked it unhealthy and our remote-command tool couldn't reach it — but nothing had crashed and nothing had been killed for using too much memory. The server was thrashing: with no swap space and containers allowed to use more memory than the machine had, Linux spent its time evicting and re-reading the files it needed to run, and made almost no progress. Adding a 2 GB swap file on every server fixed it. The more interesting lesson was about health checks.

What "hung" looked like

  • The load balancer's health checks timed out on that server.
  • Commands sent through AWS Systems Manager sat in "Pending" and never ran.
  • There was no out-of-memory kill in the logs. No process died.

A week earlier we had moved production to smaller ARM servers with 2 GB of RAM each. Our staging server — smaller still — had frozen the same way while unpacking a container image during a deploy.

Why no memory error?

Linux only kills a process when it truly runs out of memory it can reclaim. Before that point, it reclaims memory from the page cache — the copies of files it keeps in RAM, including the code of every running program.

On our servers, the containers' combined memory limits added up to about 3.3 GB on a machine with about 1.8 GB usable. Normally they don't all use their limit at once. When they did, the kernel kept evicting program code from the cache to make room, and then had to read it back from disk the moment that code ran again. That cycle — thrashing — can make a machine almost completely unresponsive while never technically running out of memory. So nothing gets killed, and nothing recovers.

Without swap, the kernel has nowhere to put memory that is allocated but rarely used, so it has to take it from the cache instead.

The fix

The server bootstrap script now creates a 2 GB swap file with a low swappiness setting (10) every time a server boots or is repaired. Low swappiness means the kernel still prefers to keep working memory in RAM, but rarely touched pages can move to disk instead of forcing program code out of the cache. We applied it live to both production servers and staging, and the deploy pipeline republishes the bootstrap script on every production deploy so a replaced server gets it too.

The part that worried us more: the unhealthy server stayed

Our server group used EC2 health checks — "is the machine running?" — not load balancer health checks — "is it answering requests?". So a server the load balancer had given up on was still healthy as far as the group was concerned. It would never be replaced. Production would quietly run on one server until someone noticed and rebooted the other.

Switching the group to load-balancer health checks sounds like the obvious fix, but it isn't straightforward for us: during a deploy, servers deliberately answer their health endpoint with an error so the balancer drains them (see how we got deploys to zero 502s). A group that replaces servers on that signal would replace them in the middle of every deploy.

So for now we added a step to our deploy checklist: after every deploy, check each target's health on the load balancer, not only whether the deploy job passed. If commands to a server sit in "Pending", it is thrashing — reboot it; it's already out of rotation.

What we took from it

  • "Not crashing" isn't "healthy". Thrashing produces no error, only slowness so extreme that the machine stops answering.
  • Small servers need swap. Even a little, with low swappiness, gives the kernel somewhere to put cold memory.
  • Memory limits that add up to more than the machine are a bet. Sometimes you lose it. Know when you've made it.
  • Know which health check your autoscaling trusts. A health check that only asks "is the VM on?" won't replace a server that has stopped working.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.

Start building free · Templates · Docs