How One Heavy Chat Could Take Our API Down
On September 28 both of our API servers ran out of memory at once. The first fix was more memory; the real one was loading chat history by size, not count.
Mythex Team · · 5 min read
On September 28, about 17 minutes after a release, both of our production API servers ran out of memory at the same moment, restarted, and for roughly five minutes the load balancer had no healthy server to send traffic to. Our first fix gave the API twice the memory. The real cause turned up later that day: opening one particular chat asked the API to load 53 MB of messages in a single request. Chat history now loads in pages capped by size as well as by count, and that chat opens with a 5 MB first page instead of a crash.
What went down
Our API runs as a container on each of two servers behind a load balancer. Each container had a memory limit of 384 MiB. At rest it uses 130–140 MiB, so the limit looked like plenty of headroom.
Both containers hit that limit together and crashed with "JavaScript heap out of memory". The heap is the memory a Node.js process uses for the objects it's working with. Because both servers failed at once, there was nowhere left to route requests. That's the worst kind of outage for a setup with two servers: redundancy only helps when the failures don't happen at the same time.
The first fix: more room
The timing pointed at the deploy. After a release, agent turns that the deploy interrupted are resumed and save their transcripts again, and that burst of work looked like the obvious source of a memory spike.
So the first change raised the API's limit from 384 MiB to 768 MiB. We applied it live on both servers the same day, and committed it to the deploy configuration so the next deploy wouldn't quietly put 384 back. A fix made by hand on a server that the next deploy then undoes is a common way for the same outage to come back.
More memory was a reasonable safety margin. It wasn't an explanation.
The real cause: one request, 71 MB
Later that day we reproduced the crash on a copy of one chat, using production's old memory limit. The copy crashed the API on its own, with no deploy involved.
When you open a project, the web app asks the API for the chat's history. The endpoint returned the newest 30 messages, whole. For most chats, 30 messages is small. For this one, the newest 30 messages were 53 MB in the database and a 71 MB response. Several of the agent's replies were 4–10 MB each, mostly the transcripts of sub-agents: the helper agents the main agent starts to work on part of a task in parallel, whose full work is saved with the reply.
Building that one response took more heap than the API had. And every open or refresh of that project sent the request again. A single chat could take down whichever server it landed on, and because the page could be reloaded, it could reach both.
"30 messages" had never been a real limit on size. It was a limit on count that happened to be small for normal chats.
The fix: pages limited by size as well as count
A page of history is now 30 messages or about 4 MB of them, whichever comes first.
It works in two steps:
- Ask for sizes first. The API asks the database for the IDs and stored sizes of the newest messages, without their contents. That query is cheap no matter how big the messages are.
- Load only what fits. Walking from newest to oldest, it adds messages until the next one would push the page past 4 MB. Only those rows are then loaded in full.
Two rules keep this safe:
- The newest message always goes on the page, whatever its size. A page that returned nothing would leave the chat stuck, with no way to scroll back.
- Nothing is cut. No message is truncated or dropped. The rest of the chat loads through the "Scroll up for earlier messages" control that older messages already used, one page at a time.
We found one more place with the same shape. When you stop a turn, the API checks the last few replies to see whether any are empty and should be cleaned up. It did that by loading eight whole messages. On a chat of multi-megabyte replies, that's the same problem in a smaller form. It now asks the database a yes-or-no question — does this message have any content? — without loading the content itself.
How we checked it
On the copy of the chat, running with the old 384 MiB limit:
- The first page is 5 MB. Before, the same request crashed the API.
- Five browser tabs opening the chat at once peak at 154 MB of memory.
- Scrolling back through the whole chat returns all 190 messages, each exactly once, in order.
The page-picking logic also has unit tests, including one built from the real sizes of that chat's newest messages. The first page comes back with one message, says there are more, and the next page starts at the one after it.
What we took from it
- A count isn't a size limit. Any endpoint that returns "the newest N" of something that can grow without bound needs a byte budget too. The data will eventually find the gap.
- The obvious suspect isn't always the cause. The deploy was a real trigger for extra work, and more memory was worth having. But the crash reproduced without a deploy, and without reproducing it we would have stopped at the memory increase.
- Two servers fail together when they get the same request. Redundancy protects against one server dying. It doesn't protect against a request that kills any server it reaches, especially one a user will naturally retry by refreshing.
- Fix it in the config, not just on the box. The memory increase went into the deploy configuration the same day, so the next deploy kept it.
This is our second recent memory story involving the agent's output. The first was a status label that queued 690 MB of strings. If you use sub-agents on Mythex, you can read about how they work in the docs.