Blog / Engineering

Keeping Image-Heavy Agent Turns Within Memory

An agent turn that kept looking at catalogue scans ran out of memory at 952 MB. Three separate causes, each holding images too long, and how we fixed them.

Mythex Team · 2026-09-29 · 5 min read

An agent turn in production ran out of memory at 952 MB while cropping products out of six catalogue scans and looking at its crops. There wasn't one leak — there were three. Every call to the model re-sent the turn's images at full size; the library we use to call the model kept every request alive until the turn ended; and the agent framework's tracing kept every call's inputs, images included, for the whole turn. We fixed all three: images are shrunk before they're sent, and each call's images are let go as soon as the call is done.

How the agent sees images

When you build on Mythex, the agent works inside your project's cloud sandbox. It can look at images — a screenshot you attach, a mockup, a file in your project, or an image it has just produced itself — by sending them to the model along with the conversation.

A turn is everything the agent does in response to one message from you. Within a turn, the agent calls the model many times: think, run a tool, look at the result, think again. Each of those calls sends the conversation so far, and that includes the turn's images, encoded as text (base64, which makes them about a third bigger).

The turn that crashed was doing something reasonable. It took six catalogue scans, about 2.4 MB each, cropped individual products out of them, and kept reading its crops back to check them. That's many model calls, each carrying several large images.

Cause 1: every call re-sent full-size images

Each model call rebuilt its request body with the turn's images at full size — about 20 MB of base64 per call. The body is built as one string in the agent's memory, so every call allocated another ~20 MB.

Sending full size didn't even buy anything. The model scales images down to about 1,300 pixels on the long side itself, so the extra pixels only cost memory and upload time.

The fix: shrink once, send small

  • Images over 400 KB are now converted to a JPEG at most 1,568 pixels on the long side before they're sent. Smaller images go as they are.
  • The conversion uses ImageMagick, which ships in every Mythex sandbox, and runs once per file. A turn that shows the same scans on every call reuses the shrunk copy. That cache is keyed by the file's path, size and modification time, so an edited file is converted again, and it has its own size cap so it can't become the thing that fills memory.
  • The budget for images in a single model call drops from 16 MB to 6 MB.

The same six scans now cost 4.4 MB per call instead of 19.7 MB.

Cause 2: the model client never let go of a request

The size of each request was only part of it. Memory also grew call after call. On our staging environment, the same kind of turn held 369 MB of live memory even after a full garbage collection — the point where anything unreachable has been freed. Heap sampling put all of the growth in the part of the model client library that encodes request bodies.

The cause was an abort listener. When you press Stop, the turn's abort signal fires, and every request listening to it cancels. We passed that one turn-wide signal to every request. The client library adds an "abort" listener to whatever signal it's given — and never removes it.

So every request of a turn stayed reachable from the turn's signal, for as long as the turn lasted: the listener, which held the request's controller, which held the request, which held the whole request body with its images.

The fix: one signal per request

Each request now gets its own signal. It's linked to the turn's signal only while that request is running — for a streamed response, until the last chunk has been read — and unlinked right after. Stop still reaches whatever request is in flight, but the turn's signal only ever holds one listener at a time.

The tests cover all three parts of that:

  • After 12 calls, no listeners are left on the turn's signal. Before the fix, there were 12.
  • Over 15 requests of 8 MB each, memory stays flat. Before, it grew by 128 MB.
  • Pressing Stop still aborts a request in the middle of its stream.

Cause 3: tracing kept every call's inputs

A heap snapshot showed a third holder: about 6 MB kept per model call. This one came from the agent framework's event streaming. It records every model call in a turn as a child of the turn, inputs included, until the turn ends. Every image attached to every call stayed in memory with it.

The fix: blank the images once the call returns

Images are only ever attached to the messages built for a specific model call. They're never stored in the conversation's saved state. So once a call returns, it's safe to empty the image data in those messages. The trace keeps a text record of the call; the pixels go.

One detail mattered: the images have to be blanked in place. The tracer can reach the same image data through more than one reference on a message, so swapping in a new, empty list would leave the original — pixels and all — still reachable. The saved conversation state is never modified.

A test runs the agent's graph under the same event streaming with a 4 MB image on every call. Without the fix, memory grew by 77 MB over 16 calls. With it, memory stays flat.

What we took from it

  • Memory problems can come in layers. Three separate causes each held images too long, and all three had to go before the turn stayed within memory.
  • Anything you hand to a library may be kept. A signal and a trace both outlived the calls they were given to. When memory grows per call, look at what long-lived objects the call touches.
  • Don't send what the other side throws away. Pixels the model would discard anyway cost us memory on every call.
  • Test memory directly. Each fix came with a test that measures heap growth or listener counts, so a regression shows up as a number, not a crash.

This is the second memory story from the agent in a short time; the first was a progress label resent 41,000 times. If you want to know more about what the agent is and how it works, see our guide to AI agents and the chat docs.

Keep reading

  • A Failed Republish Never Deletes a Working Backend — A failed republish could delete an app's live API, or leave its broken new version running. Three fixes to how Mythex undoes a publish that goes wrong.
  • Alerts That Fire Once, Not Once per Server — A once-a-day alert reached us five times in one afternoon. Why in-memory dedupe breaks with two servers and a deploy, and how shared state fixed it.
  • Moving to Bigger Workspaces Without Making Anyone Wait — A quarter of our workspaces were stuck at 1 GB after we moved to 2 GB. Our first fix made one person wait seven minutes. What we changed, and changed again.
  • Billing Databases by What They Actually Use — We used to estimate each app's database cost, and the estimate was wrong both ways. Now every database is billed from measured usage, hour by hour.
  • Catching Apps That Crash Right After Publishing — Three publishes passed our health check while their API crashed seconds into every boot. The cause: counting crashes in a list capped at five entries.
  • Keeping Long Agent Turns Inside the Context Window — Our sub-agents could grow until the model refused them, stop early to save room, or miscount tokens in Russian. Three fixes that let long AI work finish.

Start building free · Templates · Docs