The Missing Swap File That Froze Our Servers
A production server stopped answering without crashing — no out-of-memory kill, just page-cache thrash on a box with no swap. What happened and what we changed.
Mythex Team · · 3 min read
On September 24, one of our production servers stopped answering. The load balancer marked it unhealthy and our remote-command tool couldn't reach it — but nothing had crashed and nothing had been killed for using too much memory. The server was thrashing: with no swap space and containers allowed to use more memory than the machine had, Linux spent its time evicting and re-reading the files it needed to run, and made almost no progress. Adding a 2 GB swap file on every server fixed it. The more interesting lesson was about health checks.
What "hung" looked like
- The load balancer's health checks timed out on that server.
- Commands sent through AWS Systems Manager sat in "Pending" and never ran.
- There was no out-of-memory kill in the logs. No process died.
A week earlier we had moved production to smaller ARM servers with 2 GB of RAM each. Our staging server — smaller still — had frozen the same way while unpacking a container image during a deploy.
Why no memory error?
Linux only kills a process when it truly runs out of memory it can reclaim. Before that point, it reclaims memory from the page cache — the copies of files it keeps in RAM, including the code of every running program.
On our servers, the containers' combined memory limits added up to about 3.3 GB on a machine with about 1.8 GB usable. Normally they don't all use their limit at once. When they did, the kernel kept evicting program code from the cache to make room, and then had to read it back from disk the moment that code ran again. That cycle — thrashing — can make a machine almost completely unresponsive while never technically running out of memory. So nothing gets killed, and nothing recovers.
Without swap, the kernel has nowhere to put memory that is allocated but rarely used, so it has to take it from the cache instead.
The fix
The server bootstrap script now creates a 2 GB swap file with a low swappiness setting (10) every time a server boots or is repaired. Low swappiness means the kernel still prefers to keep working memory in RAM, but rarely touched pages can move to disk instead of forcing program code out of the cache. We applied it live to both production servers and staging, and the deploy pipeline republishes the bootstrap script on every production deploy so a replaced server gets it too.
The part that worried us more: the unhealthy server stayed
Our server group used EC2 health checks — "is the machine running?" — not load balancer health checks — "is it answering requests?". So a server the load balancer had given up on was still healthy as far as the group was concerned. It would never be replaced. Production would quietly run on one server until someone noticed and rebooted the other.
Switching the group to load-balancer health checks sounds like the obvious fix, but it isn't straightforward for us: during a deploy, servers deliberately answer their health endpoint with an error so the balancer drains them (see how we got deploys to zero 502s). A group that replaces servers on that signal would replace them in the middle of every deploy.
So for now we added a step to our deploy checklist: after every deploy, check each target's health on the load balancer, not only whether the deploy job passed. If commands to a server sit in "Pending", it is thrashing — reboot it; it's already out of rotation.
What we took from it
- "Not crashing" isn't "healthy". Thrashing produces no error, only slowness so extreme that the machine stops answering.
- Small servers need swap. Even a little, with low swappiness, gives the kernel somewhere to put cold memory.
- Memory limits that add up to more than the machine are a bet. Sometimes you lose it. Know when you've made it.
- Know which health check your autoscaling trusts. A health check that only asks "is the VM on?" won't replace a server that has stopped working.