Blog / Engineering

The Out-of-Memory Crash Hiding in a Progress Label

Our AI agent ran out of memory writing one large file. The cause wasn't the model or parallel sub-agents — it was a status label resent 41,000 times.

Mythex Team · 2026-09-29 · 3 min read

One evening in late September, the AI agent in a user's workspace crashed with an out-of-memory error in the middle of a turn. The obvious suspects — too many sub-agents running at once, or a context that had grown too large — were both wrong. The real cause was a status label: while the model streamed a large shell command a few characters at a time, we re-sent the entire command so far as a progress label on every piece. One 124 KB command became about 41,000 events and roughly 690 MB of queued strings.

How the agent streams its work

When you build on Mythex, the agent runs inside your project's private cloud sandbox. As the model writes, its output streams back to your browser: text appears word by word, and when it runs a tool — say, a shell command that writes a file — you see a label for it update live, so you know what it's doing.

The model streams in small pieces. With the model we were running, those pieces were often just three characters long.

The first hypotheses were wrong

The crash looked like a scaling problem, so we started there.

  • Parallel sub-agents? The turn had used only two, each holding around 1.5 MB of state. Not enough.
  • Retained traces? Our tracing kept a run tree in memory, but it wasn't anywhere near the size of the heap at the crash.

Both were plausible stories, and both were cheap to believe. Neither reproduced the crash. The lesson we keep relearning: a hypothesis isn't a cause until it reproduces the failure.

Reproducing it for real

We built a throwaway sandbox with the same memory limits as production and a fake model that streamed a large file in realistic three-character pieces. We also made the reader on the other end of the socket slow, like a real network.

That did it. The heap climbed and the process died with the same error.

The mechanism was simple once we could watch it. While the model was writing a large command — in this case a heredoc that wrote a 124 KB file — our streaming code updated the tool's label on every piece. The label wasn't the new three characters; it was the whole command so far. So the events grew: 3 characters, then 6, then 9, … up to 124 KB, around 41,000 times. The socket couldn't send them as fast as they were produced, so they piled up in its write buffer as strings — about 690 MB of them.

The fix

  • Cap the label. A status label only needs to say what's happening. It's now limited to 200 characters, and we throttle how often we re-scan the partial command.
  • Stop quadratic concatenation. Joining streamed pieces with repeated string concatenation is O(n²) for large outputs. We moved to a chunk-based approach that appends in linear time.

Then we went looking for anything else that could grow without bound on a flood of output, and bounded it:

  • Live command output shown in chat is batched and kept to the last 8,000 characters.
  • Individual event lines are capped at 4 MB.
  • The replay buffer a browser uses to catch up after reconnecting is capped at 8 MB per project.
  • Old checkpoints are pruned as a turn goes, keeping the newest state and the anchors earlier turns need.

Testing floods honestly

Two things made this harder to test than it should have been:

  1. Real models won't cooperate. Ask a model to print a gigantic file and it will usually refuse, summarise, or write it more sensibly. To test floods you need a scripted model that does exactly what you tell it, running through the real agent code.
  2. Scripts have to behave like models. Our first scripted model reused the same response ID on every reply, so the agent framework treated each reply as a replacement for the last and ended turns early. It looked like a UI refresh bug. It wasn't — it was the test harness lying to us. We also had to enforce the model's real context limit in the script, or the memory numbers were unrealistic.

What we took from it

  • Streaming UIs multiply small mistakes. Anything you send "per chunk" is sent tens of thousands of times. Send deltas, or cap the payload.
  • Bound every buffer that crosses a slow boundary. Sockets, replay logs and checkpoints all need a ceiling.
  • Reproduce before you fix. Our first two explanations sounded right and were wrong.

If you've seen a very long file appear in the Mythex chat recently without the agent stalling, this is part of why.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Image-Heavy Agent Turns Within Memory — An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.

Start building free · Templates · Docs